From efdc7bd11eccca77d6feb35d8d8d29c7d9e8cad2 Mon Sep 17 00:00:00 2001 From: Morten Hjorth-Jensen Date: Sun, 29 Sep 2024 21:22:43 +0200 Subject: [PATCH] update week 40 --- doc/pub/week40/html/._week40-bs000.html | 153 ++++--- doc/pub/week40/html/._week40-bs001.html | 186 ++++---- doc/pub/week40/html/._week40-bs002.html | 171 ++++---- doc/pub/week40/html/._week40-bs003.html | 166 ++++--- doc/pub/week40/html/._week40-bs004.html | 175 ++++---- doc/pub/week40/html/._week40-bs005.html | 183 ++++---- doc/pub/week40/html/._week40-bs006.html | 165 +++---- doc/pub/week40/html/._week40-bs007.html | 172 ++++---- doc/pub/week40/html/._week40-bs008.html | 188 ++++---- doc/pub/week40/html/._week40-bs009.html | 174 ++++---- doc/pub/week40/html/._week40-bs010.html | 205 ++++----- doc/pub/week40/html/._week40-bs011.html | 182 ++++---- doc/pub/week40/html/._week40-bs012.html | 171 ++++---- doc/pub/week40/html/._week40-bs013.html | 182 ++++---- doc/pub/week40/html/._week40-bs014.html | 251 ++++------- doc/pub/week40/html/._week40-bs015.html | 164 +++---- doc/pub/week40/html/._week40-bs016.html | 227 ++++++---- doc/pub/week40/html/._week40-bs017.html | 251 +++++++---- doc/pub/week40/html/._week40-bs018.html | 203 ++++----- doc/pub/week40/html/._week40-bs019.html | 198 +++++---- doc/pub/week40/html/._week40-bs020.html | 189 ++++---- doc/pub/week40/html/._week40-bs021.html | 233 +++++----- doc/pub/week40/html/._week40-bs022.html | 174 +++++--- doc/pub/week40/html/._week40-bs023.html | 183 ++++---- doc/pub/week40/html/._week40-bs024.html | 283 ++++++------ doc/pub/week40/html/._week40-bs025.html | 200 ++++----- doc/pub/week40/html/._week40-bs026.html | 219 ++++------ doc/pub/week40/html/._week40-bs027.html | 238 ++++++---- doc/pub/week40/html/._week40-bs028.html | 178 ++++---- doc/pub/week40/html/._week40-bs029.html | 194 +++++---- doc/pub/week40/html/._week40-bs030.html | 218 +++++----- doc/pub/week40/html/._week40-bs031.html | 185 ++++---- doc/pub/week40/html/._week40-bs032.html | 169 ++++---- doc/pub/week40/html/._week40-bs033.html | 199 +++++---- doc/pub/week40/html/._week40-bs034.html | 184 ++++---- doc/pub/week40/html/._week40-bs035.html | 211 ++++----- doc/pub/week40/html/._week40-bs036.html | 240 ++++++----- doc/pub/week40/html/._week40-bs037.html | 195 ++++----- doc/pub/week40/html/._week40-bs038.html | 193 ++++----- doc/pub/week40/html/._week40-bs039.html | 205 ++++----- doc/pub/week40/html/._week40-bs040.html | 213 ++++----- doc/pub/week40/html/._week40-bs041.html | 218 ++++++---- doc/pub/week40/html/._week40-bs042.html | 216 +++++----- doc/pub/week40/html/._week40-bs043.html | 220 ++++++---- doc/pub/week40/html/._week40-bs044.html | 220 ++++++---- doc/pub/week40/html/._week40-bs045.html | 240 +++++++---- doc/pub/week40/html/._week40-bs046.html | 250 ++++++----- doc/pub/week40/html/._week40-bs047.html | 206 +++++---- doc/pub/week40/html/._week40-bs048.html | 169 ++++---- doc/pub/week40/html/._week40-bs049.html | 221 ++++++---- doc/pub/week40/html/._week40-bs050.html | 181 ++++---- doc/pub/week40/html/._week40-bs051.html | 172 ++++---- doc/pub/week40/html/._week40-bs052.html | 174 ++++---- doc/pub/week40/html/._week40-bs053.html | 172 ++++---- doc/pub/week40/html/._week40-bs054.html | 169 ++++---- doc/pub/week40/html/._week40-bs055.html | 214 ++++----- doc/pub/week40/html/._week40-bs056.html | 233 ++++------ doc/pub/week40/html/._week40-bs057.html | 188 ++++---- doc/pub/week40/html/._week40-bs058.html | 221 ++++++---- doc/pub/week40/html/._week40-bs059.html | 252 ++++++----- doc/pub/week40/html/._week40-bs060.html | 200 +++++---- doc/pub/week40/html/._week40-bs061.html | 168 ++++---- doc/pub/week40/html/._week40-bs062.html | 195 +++++---- doc/pub/week40/html/._week40-bs063.html | 200 ++++----- doc/pub/week40/html/._week40-bs064.html | 172 ++++---- doc/pub/week40/html/._week40-bs065.html | 184 ++++---- doc/pub/week40/html/week40-bs.html | 153 ++++--- doc/pub/week40/html/week40-reveal.html | 86 ++-- doc/pub/week40/html/week40-solarized.html | 76 ++-- doc/pub/week40/html/week40.html | 76 ++-- doc/pub/week40/ipynb/ipynb-week40-src.tar.gz | Bin 34326 -> 34326 bytes doc/pub/week40/ipynb/week40.ipynb | 431 ++++++++++--------- doc/src/week40/week40.do.txt | 47 +- 73 files changed, 7517 insertions(+), 6477 deletions(-) diff --git a/doc/pub/week40/html/._week40-bs000.html b/doc/pub/week40/html/._week40-bs000.html index 5bc847095..70578a21e 100644 --- a/doc/pub/week40/html/._week40-bs000.html +++ b/doc/pub/week40/html/._week40-bs000.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -337,7 +352,7 @@ MathJax.Hub.Config({
    -

    October 2-6, 2023

    +

    September 30-October 4, 2024


    @@ -362,7 +377,7 @@ MathJax.Hub.Config({
  • 9
  • 10
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • @@ -376,7 +391,7 @@ MathJax.Hub.Config({ -->
    - © 1999-2023, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license + © 1999-2024, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license
    diff --git a/doc/pub/week40/html/._week40-bs001.html b/doc/pub/week40/html/._week40-bs001.html index 65a59cfb5..ac72280fc 100644 --- a/doc/pub/week40/html/._week40-bs001.html +++ b/doc/pub/week40/html/._week40-bs001.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -321,43 +336,6 @@ MathJax.Hub.Config({

    Plans for week 40

    -
    -
    - - -
    -
    - - -
    -
    - - -
    -
    - -

    diff --git a/doc/pub/week40/html/._week40-bs002.html b/doc/pub/week40/html/._week40-bs002.html index 5b9bfb8a4..66f4c75bb 100644 --- a/doc/pub/week40/html/._week40-bs002.html +++ b/doc/pub/week40/html/._week40-bs002.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,16 +334,20 @@ MathJax.Hub.Config({

     

     

     

    -

    Summary from last week, using gradient descent methods, limitations

    +

    Lecture Monday September 30, 2024

    +
    +
    + +
      +
    1. Stochastic Gradient descent with examples and automatic differentiation
    2. +
    3. If we get time, we start with the basics of Neural Networks, setting up the basic steps, from the simple perceptron model to the multi-layer perceptron model + +
    4. +
    +
    +
    + -

    diff --git a/doc/pub/week40/html/._week40-bs003.html b/doc/pub/week40/html/._week40-bs003.html index 93afec08b..7a7e13295 100644 --- a/doc/pub/week40/html/._week40-bs003.html +++ b/doc/pub/week40/html/._week40-bs003.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,9 +334,22 @@ MathJax.Hub.Config({

     

     

     

    -

    Overview video on Stochastic Gradient Descent

    +

    Suggested readings and videos

    +
    +
    + +
      +
    1. The lecture notes for week 40 (these notes)
    2. +
    3. For a good discussion on gradient methods, we would like to recommend Goodfellow et al section 4.3-4.5 and sections 8.3-8.6. We will come back to the latter chapter in our discussion of Neural networks as well.
    4. +
    5. For neural networks we recommend Goodfellow et al chapter 6 and Raschka et al pages 48-60
    6. +
    7. Video on gradient descent at https://www.youtube.com/watch?v=sDv4f4s2SB8
    8. +
    9. Video on stochastic gradient descent at https://www.youtube.com/watch?v=vMh0zPT0tLI
    10. +
    11. Neural Networks demystified at https://www.youtube.com/watch?v=bxe2T-V8XRs&list=PLiaHhY2iBX9hdHaRr6b7XevZtgZRa1PoU&ab_channel=WelchLabs
    12. +
    13. Building Neural Networks from scratch at URL:https://www.youtube.com/watch?v=Wo5dMEP_BbI&list=PLQVvvaa0QuDcjD5BAw2DxE6OF2tius3V3&ab_channel=sentdex"
    14. +
    +
    +
    -What is Stochastic Gradient Descent

    @@ -341,7 +369,7 @@ MathJax.Hub.Config({

  • 12
  • 13
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs004.html b/doc/pub/week40/html/._week40-bs004.html index 12cea2617..906b17dc9 100644 --- a/doc/pub/week40/html/._week40-bs004.html +++ b/doc/pub/week40/html/._week40-bs004.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,19 +334,19 @@ MathJax.Hub.Config({

     

     

     

    -

    Batches and mini-batches

    - -

    In gradient descent we compute the cost function and its gradient for all data points we have.

    - -

    In large-scale applications such as the ILSVRC challenge, the -training data can have on order of millions of examples. Hence, it -seems wasteful to compute the full cost function over the entire -training set in order to perform only a single parameter update. A -very common approach to addressing this challenge is to compute the -gradient over batches of the training data. For example, a typical batch could contain some thousand examples from -an entire training set of several millions. This batch is then used to -perform a parameter update. -

    +

    Lab sessions Tuesday and Wednesday

    +
    +
    + + +
    +
    +

    @@ -352,7 +367,7 @@ perform a parameter update.

  • 13
  • 14
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs005.html b/doc/pub/week40/html/._week40-bs005.html index b1fba64b4..81314ba65 100644 --- a/doc/pub/week40/html/._week40-bs005.html +++ b/doc/pub/week40/html/._week40-bs005.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,32 +334,16 @@ MathJax.Hub.Config({

     

     

     

    -

    Stochastic Gradient Descent (SGD)

    - -

    In stochastic gradient descent, the extreme case is the case where we -have only one batch, that is we include the whole data set. -

    - -

    This process is called Stochastic Gradient -Descent (SGD) (or also sometimes on-line gradient descent). This is -relatively less common to see because in practice due to vectorized -code optimizations it can be computationally much more efficient to -evaluate the gradient for 100 examples, than the gradient for one -example 100 times. Even though SGD technically refers to using a -single example at a time to evaluate the gradient, you will hear -people use the term SGD even when referring to mini-batch gradient -descent (i.e. mentions of MGD for “Minibatch Gradient Descent”, or BGD -for “Batch gradient descent” are rare to see), where it is usually -assumed that mini-batches are used. The size of the mini-batch is a -hyperparameter but it is not very common to cross-validate or bootstrap it. It is -usually based on memory constraints (if any), or set to some value, -e.g. 32, 64 or 128. We use powers of 2 in practice because many -vectorized operation implementations work faster when their inputs are -sized in powers of 2. -

    - -

    In our notes with SGD we mean stochastic gradient descent with mini-batches.

    +

    Summary from last week, using gradient descent methods, limitations

    +

    diff --git a/doc/pub/week40/html/._week40-bs006.html b/doc/pub/week40/html/._week40-bs006.html index 8a678db09..47776062a 100644 --- a/doc/pub/week40/html/._week40-bs006.html +++ b/doc/pub/week40/html/._week40-bs006.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,21 +334,9 @@ MathJax.Hub.Config({

     

     

     

    -

    Stochastic Gradient Descent

    - -

    Stochastic gradient descent (SGD) and variants thereof address some of -the shortcomings of the Gradient descent method discussed above. -

    - -

    The underlying idea of SGD comes from the observation that the cost -function, which we want to minimize, can almost always be written as a -sum over \( n \) data points \( \{\mathbf{x}_i\}_{i=1}^n \), -

    -$$ -C(\mathbf{\beta}) = \sum_{i=1}^n c_i(\mathbf{x}_i, -\mathbf{\beta}). -$$ +

    Overview video on Stochastic Gradient Descent

    +What is Stochastic Gradient Descent

    @@ -356,7 +359,7 @@ $$

  • 15
  • 16
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs007.html b/doc/pub/week40/html/._week40-bs007.html index 777c770ff..aa127487b 100644 --- a/doc/pub/week40/html/._week40-bs007.html +++ b/doc/pub/week40/html/._week40-bs007.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,21 +334,18 @@ MathJax.Hub.Config({

     

     

     

    -

    Computation of gradients

    +

    Batches and mini-batches

    -

    This in turn means that the gradient can be -computed as a sum over \( i \)-gradients -

    -$$ -\nabla_\beta C(\mathbf{\beta}) = \sum_i^n \nabla_\beta c_i(\mathbf{x}_i, -\mathbf{\beta}). -$$ +

    In gradient descent we compute the cost function and its gradient for all data points we have.

    -

    Stochasticity/randomness is introduced by only taking the -gradient on a subset of the data called minibatches. If there are \( n \) -data points and the size of each minibatch is \( M \), there will be \( n/M \) -minibatches. We denote these minibatches by \( B_k \) where -\( k=1,\cdots,n/M \). +

    In large-scale applications such as the ILSVRC challenge, the +training data can have on order of millions of examples. Hence, it +seems wasteful to compute the full cost function over the entire +training set in order to perform only a single parameter update. A +very common approach to addressing this challenge is to compute the +gradient over batches of the training data. For example, a typical batch could contain some thousand examples from +an entire training set of several millions. This batch is then used to +perform a parameter update.

    @@ -358,7 +370,7 @@ minibatches. We denote these minibatches by \( B_k \) where

  • 16
  • 17
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs008.html b/doc/pub/week40/html/._week40-bs008.html index b9f8ef17b..a1038fcb6 100644 --- a/doc/pub/week40/html/._week40-bs008.html +++ b/doc/pub/week40/html/._week40-bs008.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,28 +334,31 @@ MathJax.Hub.Config({

     

     

     

    -

    SGD example

    -

    As an example, suppose we have \( 10 \) data points \( (\mathbf{x}_1,\cdots, \mathbf{x}_{10}) \) -and we choose to have \( M=5 \) minibathces, -then each minibatch contains two data points. In particular we have -\( B_1 = (\mathbf{x}_1,\mathbf{x}_2), \cdots, B_5 = -(\mathbf{x}_9,\mathbf{x}_{10}) \). Note that if you choose \( M=1 \) you -have only a single batch with all data points and on the other extreme, -you may choose \( M=n \) resulting in a minibatch for each datapoint, i.e -\( B_k = \mathbf{x}_k \). +

    Stochastic Gradient Descent (SGD)

    + +

    In stochastic gradient descent, the extreme case is the case where we +have only one batch, that is we include the whole data set.

    -

    The idea is now to approximate the gradient by replacing the sum over -all data points with a sum over the data points in one the minibatches -picked at random in each gradient descent step +

    This process is called Stochastic Gradient +Descent (SGD) (or also sometimes on-line gradient descent). This is +relatively less common to see because in practice due to vectorized +code optimizations it can be computationally much more efficient to +evaluate the gradient for 100 examples, than the gradient for one +example 100 times. Even though SGD technically refers to using a +single example at a time to evaluate the gradient, you will hear +people use the term SGD even when referring to mini-batch gradient +descent (i.e. mentions of MGD for “Minibatch Gradient Descent”, or BGD +for “Batch gradient descent” are rare to see), where it is usually +assumed that mini-batches are used. The size of the mini-batch is a +hyperparameter but it is not very common to cross-validate or bootstrap it. It is +usually based on memory constraints (if any), or set to some value, +e.g. 32, 64 or 128. We use powers of 2 in practice because many +vectorized operation implementations work faster when their inputs are +sized in powers of 2.

    -$$ -\nabla_{\beta} -C(\mathbf{\beta}) = \sum_{i=1}^n \nabla_\beta c_i(\mathbf{x}_i, -\mathbf{\beta}) \rightarrow \sum_{i \in B_k}^n \nabla_\beta -c_i(\mathbf{x}_i, \mathbf{\beta}). -$$ +

    In our notes with SGD we mean stochastic gradient descent with mini-batches.

    @@ -365,7 +383,7 @@ $$

  • 17
  • 18
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs009.html b/doc/pub/week40/html/._week40-bs009.html index f4d23cae6..1245c434d 100644 --- a/doc/pub/week40/html/._week40-bs009.html +++ b/doc/pub/week40/html/._week40-bs009.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,21 +334,22 @@ MathJax.Hub.Config({

     

     

     

    -

    The gradient step

    +

    Stochastic Gradient Descent

    -

    Thus a gradient descent step now looks like

    -$$ -\beta_{j+1} = \beta_j - \gamma_j \sum_{i \in B_k}^n \nabla_\beta c_i(\mathbf{x}_i, -\mathbf{\beta}) -$$ - -

    where \( k \) is picked at random with equal -probability from \( [1,n/M] \). An iteration over the number of -minibathces (n/M) is commonly referred to as an epoch. Thus it is -typical to choose a number of epochs and for each epoch iterate over -the number of minibatches, as exemplified in the code below. +

    Stochastic gradient descent (SGD) and variants thereof address some of +the shortcomings of the Gradient descent method discussed above.

    +

    The underlying idea of SGD comes from the observation that the cost +function, which we want to minimize, can almost always be written as a +sum over \( n \) data points \( \{\mathbf{x}_i\}_{i=1}^n \), +

    +$$ +C(\mathbf{\beta}) = \sum_{i=1}^n c_i(\mathbf{x}_i, +\mathbf{\beta}). +$$ + +

    diff --git a/doc/pub/week40/html/._week40-bs010.html b/doc/pub/week40/html/._week40-bs010.html index 6771de428..0e94a97ba 100644 --- a/doc/pub/week40/html/._week40-bs010.html +++ b/doc/pub/week40/html/._week40-bs010.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,51 +334,21 @@ MathJax.Hub.Config({

     

     

     

    -

    Simple example code

    +

    Computation of gradients

    +

    This in turn means that the gradient can be +computed as a sum over \( i \)-gradients +

    +$$ +\nabla_\beta C(\mathbf{\beta}) = \sum_i^n \nabla_\beta c_i(\mathbf{x}_i, +\mathbf{\beta}). +$$ - -
    -
    -
    -
    -
    -
    import numpy as np 
    -
    -n = 100 #100 datapoints 
    -M = 5   #size of each minibatch
    -m = int(n/M) #number of minibatches
    -n_epochs = 10 #number of epochs
    -
    -j = 0
    -for epoch in range(1,n_epochs+1):
    -    for i in range(m):
    -        k = np.random.randint(m) #Pick the k-th minibatch at random
    -        #Compute the gradient using the data in minibatch Bk
    -        #Compute new suggestion for 
    -        j += 1
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - -

    Taking the gradient only on a subset of the data has two important -benefits. First, it introduces randomness which decreases the chance -that our opmization scheme gets stuck in a local minima. Second, if -the size of the minibatches are small relative to the number of -datapoints (\( M < n \)), the computation of the gradient is much -cheaper since we sum over the datapoints in the \( k-th \) minibatch and not -all \( n \) datapoints. +

    Stochasticity/randomness is introduced by only taking the +gradient on a subset of the data called minibatches. If there are \( n \) +data points and the size of each minibatch is \( M \), there will be \( n/M \) +minibatches. We denote these minibatches by \( B_k \) where +\( k=1,\cdots,n/M \).

    @@ -391,7 +376,7 @@ all \( n \) datapoints.

  • 19
  • 20
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs011.html b/doc/pub/week40/html/._week40-bs011.html index 9f1d0deef..5c182992f 100644 --- a/doc/pub/week40/html/._week40-bs011.html +++ b/doc/pub/week40/html/._week40-bs011.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,20 +334,29 @@ MathJax.Hub.Config({

     

     

     

    -

    When do we stop?

    - -

    A natural question is when do we stop the search for a new minimum? -One possibility is to compute the full gradient after a given number -of epochs and check if the norm of the gradient is smaller than some -threshold and stop if true. However, the condition that the gradient -is zero is valid also for local minima, so this would only tell us -that we are close to a local/global minimum. However, we could also -evaluate the cost function at this point, store the result and -continue the search. If the test kicks in at a later stage we can -compare the values of the cost function and keep the \( \beta \) that -gave the lowest value. +

    SGD example

    +

    As an example, suppose we have \( 10 \) data points \( (\mathbf{x}_1,\cdots, \mathbf{x}_{10}) \) +and we choose to have \( M=5 \) minibathces, +then each minibatch contains two data points. In particular we have +\( B_1 = (\mathbf{x}_1,\mathbf{x}_2), \cdots, B_5 = +(\mathbf{x}_9,\mathbf{x}_{10}) \). Note that if you choose \( M=1 \) you +have only a single batch with all data points and on the other extreme, +you may choose \( M=n \) resulting in a minibatch for each datapoint, i.e +\( B_k = \mathbf{x}_k \).

    +

    The idea is now to approximate the gradient by replacing the sum over +all data points with a sum over the data points in one the minibatches +picked at random in each gradient descent step +

    +$$ +\nabla_{\beta} +C(\mathbf{\beta}) = \sum_{i=1}^n \nabla_\beta c_i(\mathbf{x}_i, +\mathbf{\beta}) \rightarrow \sum_{i \in B_k}^n \nabla_\beta +c_i(\mathbf{x}_i, \mathbf{\beta}). +$$ + +

    diff --git a/doc/pub/week40/html/._week40-bs012.html b/doc/pub/week40/html/._week40-bs012.html index 6db89d90c..924de2441 100644 --- a/doc/pub/week40/html/._week40-bs012.html +++ b/doc/pub/week40/html/._week40-bs012.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,17 +334,19 @@ MathJax.Hub.Config({

     

     

     

    -

    Slightly different approach

    +

    The gradient step

    -

    Another approach is to let the step length \( \gamma_j \) depend on the -number of epochs in such a way that it becomes very small after a -reasonable time such that we do not move at all. Such approaches are -also called scaling. There are many such ways to scale the learning -rate -and discussions here. See -also -https://towardsdatascience.com/learning-rate-schedules-and-adaptive-learning-rate-methods-for-deep-learning-2c8f433990d1 -for a discussion of different scaling functions for the learning rate. +

    Thus a gradient descent step now looks like

    +$$ +\beta_{j+1} = \beta_j - \gamma_j \sum_{i \in B_k}^n \nabla_\beta c_i(\mathbf{x}_i, +\mathbf{\beta}) +$$ + +

    where \( k \) is picked at random with equal +probability from \( [1,n/M] \). An iteration over the number of +minibathces (n/M) is commonly referred to as an epoch. Thus it is +typical to choose a number of epochs and for each epoch iterate over +the number of minibatches, as exemplified in the code below.

    @@ -357,7 +374,7 @@ for a discussion of different scaling functions for the learning rate.

  • 21
  • 22
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs013.html b/doc/pub/week40/html/._week40-bs013.html index cd7a53a73..698ed48ff 100644 --- a/doc/pub/week40/html/._week40-bs013.html +++ b/doc/pub/week40/html/._week40-bs013.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,16 +334,7 @@ MathJax.Hub.Config({

     

     

     

    -

    Time decay rate

    - -

    As an example, let \( e = 0,1,2,3,\cdots \) denote the current epoch and let \( t_0, t_1 > 0 \) be two fixed numbers. Furthermore, let \( t = e \cdot m + i \) where \( m \) is the number of minibatches and \( i=0,\cdots,m-1 \). Then the function $$\gamma_j(t; t_0, t_1) = \frac{t_0}{t+t_1} $$ goes to zero as the number of epochs gets large. I.e. we start with a step length \( \gamma_j (0; t_0, t_1) = t_0/t_1 \) which decays in time \( t \).

    - -

    In this way we can fix the number of epochs, compute \( \beta \) and -evaluate the cost function at the end. Repeating the computation will -give a different result since the scheme is random by design. Then we -pick the final \( \beta \) that gives the lowest value of the cost -function. -

    +

    Simple example code

    @@ -339,28 +345,18 @@ function.
    import numpy as np 
     
    -def step_length(t,t0,t1):
    -    return t0/(t+t1)
    -
     n = 100 #100 datapoints 
     M = 5   #size of each minibatch
     m = int(n/M) #number of minibatches
    -n_epochs = 500 #number of epochs
    -t0 = 1.0
    -t1 = 10
    +n_epochs = 10 #number of epochs
     
    -gamma_j = t0/t1
     j = 0
     for epoch in range(1,n_epochs+1):
         for i in range(m):
             k = np.random.randint(m) #Pick the k-th minibatch at random
             #Compute the gradient using the data in minibatch Bk
    -        #Compute new suggestion for beta
    -        t = epoch*m+i
    -        gamma_j = step_length(t,t0,t1)
    +        #Compute new suggestion for 
             j += 1
    -
    -print("gamma_j after %d epochs: %g" % (n_epochs,gamma_j))
     
    @@ -376,6 +372,14 @@ j = 0 +

    Taking the gradient only on a subset of the data has two important +benefits. First, it introduces randomness which decreases the chance +that our opmization scheme gets stuck in a local minima. Second, if +the size of the minibatches are small relative to the number of +datapoints (\( M < n \)), the computation of the gradient is much +cheaper since we sum over the datapoints in the \( k-th \) minibatch and not +all \( n \) datapoints. +

    @@ -402,7 +406,7 @@ j = 0

  • 22
  • 23
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs014.html b/doc/pub/week40/html/._week40-bs014.html index 833ff1fed..39b79c372 100644 --- a/doc/pub/week40/html/._week40-bs014.html +++ b/doc/pub/week40/html/._week40-bs014.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,97 +334,19 @@ MathJax.Hub.Config({

     

     

     

    -

    Code with a Number of Minibatches which varies

    - -

    In the code here we vary the number of mini-batches.

    - - -
    -
    -
    -
    -
    -
    # Importing various packages
    -from math import exp, sqrt
    -from random import random, seed
    -import numpy as np
    -import matplotlib.pyplot as plt
    -
    -n = 100
    -x = 2*np.random.rand(n,1)
    -y = 4+3*x+np.random.randn(n,1)
    -
    -X = np.c_[np.ones((n,1)), x]
    -XT_X = X.T @ X
    -theta_linreg = np.linalg.inv(X.T @ X) @ (X.T @ y)
    -print("Own inversion")
    -print(theta_linreg)
    -# Hessian matrix
    -H = (2.0/n)* XT_X
    -EigValues, EigVectors = np.linalg.eig(H)
    -print(f"Eigenvalues of Hessian Matrix:{EigValues}")
    -
    -theta = np.random.randn(2,1)
    -eta = 1.0/np.max(EigValues)
    -Niterations = 1000
    -
    -
    -for iter in range(Niterations):
    -    gradients = 2.0/n*X.T @ ((X @ theta)-y)
    -    theta -= eta*gradients
    -print("theta from own gd")
    -print(theta)
    -
    -xnew = np.array([[0],[2]])
    -Xnew = np.c_[np.ones((2,1)), xnew]
    -ypredict = Xnew.dot(theta)
    -ypredict2 = Xnew.dot(theta_linreg)
    -
    -n_epochs = 50
    -M = 5   #size of each minibatch
    -m = int(n/M) #number of minibatches
    -t0, t1 = 5, 50
    -
    -def learning_schedule(t):
    -    return t0/(t+t1)
    -
    -theta = np.random.randn(2,1)
    -
    -for epoch in range(n_epochs):
    -# Can you figure out a better way of setting up the contributions to each batch?
    -    for i in range(m):
    -        random_index = M*np.random.randint(m)
    -        xi = X[random_index:random_index+M]
    -        yi = y[random_index:random_index+M]
    -        gradients = (2.0/M)* xi.T @ ((xi @ theta)-yi)
    -        eta = learning_schedule(epoch*m+i)
    -        theta = theta - eta*gradients
    -print("theta from own sdg")
    -print(theta)
    -
    -plt.plot(xnew, ypredict, "r-")
    -plt.plot(xnew, ypredict2, "b-")
    -plt.plot(x, y ,'ro')
    -plt.axis([0,2.0,0, 15.0])
    -plt.xlabel(r'$x$')
    -plt.ylabel(r'$y$')
    -plt.title(r'Random numbers ')
    -plt.show()
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    +

    When do we stop?

    +

    A natural question is when do we stop the search for a new minimum? +One possibility is to compute the full gradient after a given number +of epochs and check if the norm of the gradient is smaller than some +threshold and stop if true. However, the condition that the gradient +is zero is valid also for local minima, so this would only tell us +that we are close to a local/global minimum. However, we could also +evaluate the cost function at this point, store the result and +continue the search. If the test kicks in at a later stage we can +compare the values of the cost function and keep the \( \beta \) that +gave the lowest value. +

    @@ -436,7 +373,7 @@ plt.show()

  • 23
  • 24
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs015.html b/doc/pub/week40/html/._week40-bs015.html index f10af4e22..fa80fee3c 100644 --- a/doc/pub/week40/html/._week40-bs015.html +++ b/doc/pub/week40/html/._week40-bs015.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,12 +334,17 @@ MathJax.Hub.Config({

     

     

     

    -

    Replace or not

    +

    Slightly different approach

    -

    In the above code, we have use replacement in setting up the -mini-batches. The discussion -here may be -useful. +

    Another approach is to let the step length \( \gamma_j \) depend on the +number of epochs in such a way that it becomes very small after a +reasonable time such that we do not move at all. Such approaches are +also called scaling. There are many such ways to scale the learning +rate +and discussions here. See +also +https://towardsdatascience.com/learning-rate-schedules-and-adaptive-learning-rate-methods-for-deep-learning-2c8f433990d1 +for a discussion of different scaling functions for the learning rate.

    @@ -352,7 +372,7 @@ useful.

  • 24
  • 25
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs016.html b/doc/pub/week40/html/._week40-bs016.html index 15db620ae..aada1dd4a 100644 --- a/doc/pub/week40/html/._week40-bs016.html +++ b/doc/pub/week40/html/._week40-bs016.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,39 +334,63 @@ MathJax.Hub.Config({

     

     

     

    -

    Momentum based GD

    +

    Time decay rate

    -

    The stochastic gradient descent (SGD) is almost always used with a -momentum or inertia term that serves as a memory of the direction we -are moving in parameter space. This is typically implemented as -follows +

    As an example, let \( e = 0,1,2,3,\cdots \) denote the current epoch and let \( t_0, t_1 > 0 \) be two fixed numbers. Furthermore, let \( t = e \cdot m + i \) where \( m \) is the number of minibatches and \( i=0,\cdots,m-1 \). Then the function $$\gamma_j(t; t_0, t_1) = \frac{t_0}{t+t_1} $$ goes to zero as the number of epochs gets large. I.e. we start with a step length \( \gamma_j (0; t_0, t_1) = t_0/t_1 \) which decays in time \( t \).

    + +

    In this way we can fix the number of epochs, compute \( \beta \) and +evaluate the cost function at the end. Repeating the computation will +give a different result since the scheme is random by design. Then we +pick the final \( \beta \) that gives the lowest value of the cost +function.

    -$$ -\begin{align} -\mathbf{v}_{t}&=\gamma \mathbf{v}_{t-1}+\eta_{t}\nabla_\theta E(\boldsymbol{\theta}_t) \nonumber \\ -\boldsymbol{\theta}_{t+1}&= \boldsymbol{\theta}_t -\mathbf{v}_{t}, -\tag{1} -\end{align} -$$ -

    where we have introduced a momentum parameter \( \gamma \), with -\( 0\le\gamma\le 1 \), and for brevity we dropped the explicit notation to -indicate the gradient is to be taken over a different mini-batch at -each step. We call this algorithm gradient descent with momentum -(GDM). From these equations, it is clear that \( \mathbf{v}_t \) is a -running average of recently encountered gradients and -\( (1-\gamma)^{-1} \) sets the characteristic time scale for the memory -used in the averaging procedure. Consistent with this, when -\( \gamma=0 \), this just reduces down to ordinary SGD as discussed -earlier. An equivalent way of writing the updates is -

    + +
    +
    +
    +
    +
    +
    import numpy as np 
     
    -$$
    -\Delta \boldsymbol{\theta}_{t+1} = \gamma \Delta \boldsymbol{\theta}_t -\ \eta_{t}\nabla_\theta E(\boldsymbol{\theta}_t),
    -$$
    +def step_length(t,t0,t1):
    +    return t0/(t+t1)
    +
    +n = 100 #100 datapoints 
    +M = 5   #size of each minibatch
    +m = int(n/M) #number of minibatches
    +n_epochs = 500 #number of epochs
    +t0 = 1.0
    +t1 = 10
    +
    +gamma_j = t0/t1
    +j = 0
    +for epoch in range(1,n_epochs+1):
    +    for i in range(m):
    +        k = np.random.randint(m) #Pick the k-th minibatch at random
    +        #Compute the gradient using the data in minibatch Bk
    +        #Compute new suggestion for beta
    +        t = epoch*m+i
    +        gamma_j = step_length(t,t0,t1)
    +        j += 1
    +
    +print("gamma_j after %d epochs: %g" % (n_epochs,gamma_j))
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    -

    where we have defined \( \Delta \boldsymbol{\theta}_{t}= \boldsymbol{\theta}_t-\boldsymbol{\theta}_{t-1} \).

    @@ -378,7 +417,7 @@ $$

  • 25
  • 26
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs017.html b/doc/pub/week40/html/._week40-bs017.html index be2351601..b6b47b8b3 100644 --- a/doc/pub/week40/html/._week40-bs017.html +++ b/doc/pub/week40/html/._week40-bs017.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,30 +334,96 @@ MathJax.Hub.Config({

     

     

     

    -

    More on momentum based approaches

    +

    Code with a Number of Minibatches which varies

    -

    Let us try to get more intuition from these equations. It is helpful -to consider a simple physical analogy with a particle of mass \( m \) -moving in a viscous medium with drag coefficient \( \mu \) and potential -\( E(\mathbf{w}) \). If we denote the particle's position by \( \mathbf{w} \), -then its motion is described by -

    +

    In the code here we vary the number of mini-batches.

    -$$ -m {d^2 \mathbf{w} \over dt^2} + \mu {d \mathbf{w} \over dt }= -\nabla_w E(\mathbf{w}). -$$ + +
    +
    +
    +
    +
    +
    # Importing various packages
    +from math import exp, sqrt
    +from random import random, seed
    +import numpy as np
    +import matplotlib.pyplot as plt
     
    -

    We can discretize this equation in the usual way to get

    +n = 100 +x = 2*np.random.rand(n,1) +y = 4+3*x+np.random.randn(n,1) -$$ -m { \mathbf{w}_{t+\Delta t}-2 \mathbf{w}_{t} +\mathbf{w}_{t-\Delta t} \over (\Delta t)^2}+\mu {\mathbf{w}_{t+\Delta t}- \mathbf{w}_{t} \over \Delta t} = -\nabla_w E(\mathbf{w}). -$$ +X = np.c_[np.ones((n,1)), x] +XT_X = X.T @ X +theta_linreg = np.linalg.inv(X.T @ X) @ (X.T @ y) +print("Own inversion") +print(theta_linreg) +# Hessian matrix +H = (2.0/n)* XT_X +EigValues, EigVectors = np.linalg.eig(H) +print(f"Eigenvalues of Hessian Matrix:{EigValues}") -

    Rearranging this equation, we can rewrite this as

    +theta = np.random.randn(2,1) +eta = 1.0/np.max(EigValues) +Niterations = 1000 -$$ -\Delta \mathbf{w}_{t +\Delta t}= - { (\Delta t)^2 \over m +\mu \Delta t} \nabla_w E(\mathbf{w})+ {m \over m +\mu \Delta t} \Delta \mathbf{w}_t. -$$ + +for iter in range(Niterations): + gradients = 2.0/n*X.T @ ((X @ theta)-y) + theta -= eta*gradients +print("theta from own gd") +print(theta) + +xnew = np.array([[0],[2]]) +Xnew = np.c_[np.ones((2,1)), xnew] +ypredict = Xnew.dot(theta) +ypredict2 = Xnew.dot(theta_linreg) + +n_epochs = 50 +M = 5 #size of each minibatch +m = int(n/M) #number of minibatches +t0, t1 = 5, 50 + +def learning_schedule(t): + return t0/(t+t1) + +theta = np.random.randn(2,1) + +for epoch in range(n_epochs): +# Can you figure out a better way of setting up the contributions to each batch? + for i in range(m): + random_index = M*np.random.randint(m) + xi = X[random_index:random_index+M] + yi = y[random_index:random_index+M] + gradients = (2.0/M)* xi.T @ ((xi @ theta)-yi) + eta = learning_schedule(epoch*m+i) + theta = theta - eta*gradients +print("theta from own sdg") +print(theta) + +plt.plot(xnew, ypredict, "r-") +plt.plot(xnew, ypredict2, "b-") +plt.plot(x, y ,'ro') +plt.axis([0,2.0,0, 15.0]) +plt.xlabel(r'$x$') +plt.ylabel(r'$y$') +plt.title(r'Random numbers ') +plt.show() +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +

    @@ -370,7 +451,7 @@ $$

  • 26
  • 27
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs018.html b/doc/pub/week40/html/._week40-bs018.html index d53a8fe4c..f74b539d9 100644 --- a/doc/pub/week40/html/._week40-bs018.html +++ b/doc/pub/week40/html/._week40-bs018.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,58 +334,14 @@ MathJax.Hub.Config({

     

     

     

    -

    Momentum parameter

    +

    Replace or not

    -

    Notice that this equation is identical to previous one if we identify -the position of the particle, \( \mathbf{w} \), with the parameters -\( \boldsymbol{\theta} \). This allows us to identify the momentum -parameter and learning rate with the mass of the particle and the -viscous drag as: +

    In the above code, we have use replacement in setting up the +mini-batches. The discussion +here may be +useful.

    -$$ -\gamma= {m \over m +\mu \Delta t }, \qquad \eta = {(\Delta t)^2 \over m +\mu \Delta t}. -$$ - -

    Thus, as the name suggests, the momentum parameter is proportional to -the mass of the particle and effectively provides inertia. -Furthermore, in the large viscosity/small learning rate limit, our -memory time scales as \( (1-\gamma)^{-1} \approx m/(\mu \Delta t) \). -

    - -

    Why is momentum useful? SGD momentum helps the gradient descent -algorithm gain speed in directions with persistent but small gradients -even in the presence of stochasticity, while suppressing oscillations -in high-curvature directions. This becomes especially important in -situations where the landscape is shallow and flat in some directions -and narrow and steep in others. It has been argued that first-order -methods (with appropriate initial conditions) can perform comparable -to more expensive second order methods, especially in the context of -complex deep learning models. -

    - -

    These beneficial properties of momentum can sometimes become even more -pronounced by using a slight modification of the classical momentum -algorithm called Nesterov Accelerated Gradient (NAG). -

    - -

    In the NAG algorithm, rather than calculating the gradient at the -current parameters, \( \nabla_\theta E(\boldsymbol{\theta}_t) \), one -calculates the gradient at the expected value of the parameters given -our current momentum, \( \nabla_\theta E(\boldsymbol{\theta}_t +\gamma -\mathbf{v}_{t-1}) \). This yields the NAG update rule -

    - -$$ -\begin{align} -\mathbf{v}_{t}&=\gamma \mathbf{v}_{t-1}+\eta_{t}\nabla_\theta E(\boldsymbol{\theta}_t +\gamma \mathbf{v}_{t-1}) \nonumber \\ -\boldsymbol{\theta}_{t+1}&= \boldsymbol{\theta}_t -\mathbf{v}_{t}. -\tag{2} -\end{align} -$$ - -

    One of the major advantages of NAG is that it allows for the use of a larger learning rate than GDM for the same choice of \( \gamma \).

    -

    diff --git a/doc/pub/week40/html/._week40-bs019.html b/doc/pub/week40/html/._week40-bs019.html index 0b84fc505..c5c08b2ed 100644 --- a/doc/pub/week40/html/._week40-bs019.html +++ b/doc/pub/week40/html/._week40-bs019.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,31 +334,40 @@ MathJax.Hub.Config({

     

     

     

    -

    Second moment of the gradient

    +

    Momentum based GD

    -

    In stochastic gradient descent, with and without momentum, we still -have to specify a schedule for tuning the learning rates \( \eta_t \) -as a function of time. As discussed in the context of Newton's -method, this presents a number of dilemmas. The learning rate is -limited by the steepest direction which can change depending on the -current position in the landscape. To circumvent this problem, ideally -our algorithm would keep track of curvature and take large steps in -shallow, flat directions and small steps in steep, narrow directions. -Second-order methods accomplish this by calculating or approximating -the Hessian and normalizing the learning rate by the -curvature. However, this is very computationally expensive for -extremely large models. Ideally, we would like to be able to -adaptively change the step size to match the landscape without paying -the steep computational price of calculating or approximating -Hessians. +

    The stochastic gradient descent (SGD) is almost always used with a +momentum or inertia term that serves as a memory of the direction we +are moving in parameter space. This is typically implemented as +follows

    -

    Recently, a number of methods have been introduced that accomplish -this by tracking not only the gradient, but also the second moment of -the gradient. These methods include AdaGrad, AdaDelta, Root Mean Squared Propagation (RMS-Prop), and -ADAM. +$$ +\begin{align} +\mathbf{v}_{t}&=\gamma \mathbf{v}_{t-1}+\eta_{t}\nabla_\theta E(\boldsymbol{\theta}_t) \nonumber \\ +\boldsymbol{\theta}_{t+1}&= \boldsymbol{\theta}_t -\mathbf{v}_{t}, +\tag{1} +\end{align} +$$ + +

    where we have introduced a momentum parameter \( \gamma \), with +\( 0\le\gamma\le 1 \), and for brevity we dropped the explicit notation to +indicate the gradient is to be taken over a different mini-batch at +each step. We call this algorithm gradient descent with momentum +(GDM). From these equations, it is clear that \( \mathbf{v}_t \) is a +running average of recently encountered gradients and +\( (1-\gamma)^{-1} \) sets the characteristic time scale for the memory +used in the averaging procedure. Consistent with this, when +\( \gamma=0 \), this just reduces down to ordinary SGD as discussed +earlier. An equivalent way of writing the updates is

    +$$ +\Delta \boldsymbol{\theta}_{t+1} = \gamma \Delta \boldsymbol{\theta}_t -\ \eta_{t}\nabla_\theta E(\boldsymbol{\theta}_t), +$$ + +

    where we have defined \( \Delta \boldsymbol{\theta}_{t}= \boldsymbol{\theta}_t-\boldsymbol{\theta}_{t-1} \).

    +

    diff --git a/doc/pub/week40/html/._week40-bs020.html b/doc/pub/week40/html/._week40-bs020.html index 89aebe3b3..bcfede1ba 100644 --- a/doc/pub/week40/html/._week40-bs020.html +++ b/doc/pub/week40/html/._week40-bs020.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,33 +334,31 @@ MathJax.Hub.Config({

     

     

     

    -

    RMS prop

    +

    More on momentum based approaches

    -

    In RMS prop, in addition to keeping a running average of the first -moment of the gradient, we also keep track of the second moment -denoted by \( \mathbf{s}_t=\mathbb{E}[\mathbf{g}_t^2] \). The update rule -for RMS prop is given by +

    Let us try to get more intuition from these equations. It is helpful +to consider a simple physical analogy with a particle of mass \( m \) +moving in a viscous medium with drag coefficient \( \mu \) and potential +\( E(\mathbf{w}) \). If we denote the particle's position by \( \mathbf{w} \), +then its motion is described by

    $$ -\begin{align} -\mathbf{g}_t &= \nabla_\theta E(\boldsymbol{\theta}) -\tag{3}\\ -\mathbf{s}_t &=\beta \mathbf{s}_{t-1} +(1-\beta)\mathbf{g}_t^2 \nonumber \\ -\boldsymbol{\theta}_{t+1}&=&\boldsymbol{\theta}_t - \eta_t { \mathbf{g}_t \over \sqrt{\mathbf{s}_t +\epsilon}}, \nonumber -\end{align} +m {d^2 \mathbf{w} \over dt^2} + \mu {d \mathbf{w} \over dt }= -\nabla_w E(\mathbf{w}). +$$ + +

    We can discretize this equation in the usual way to get

    + +$$ +m { \mathbf{w}_{t+\Delta t}-2 \mathbf{w}_{t} +\mathbf{w}_{t-\Delta t} \over (\Delta t)^2}+\mu {\mathbf{w}_{t+\Delta t}- \mathbf{w}_{t} \over \Delta t} = -\nabla_w E(\mathbf{w}). +$$ + +

    Rearranging this equation, we can rewrite this as

    + +$$ +\Delta \mathbf{w}_{t +\Delta t}= - { (\Delta t)^2 \over m +\mu \Delta t} \nabla_w E(\mathbf{w})+ {m \over m +\mu \Delta t} \Delta \mathbf{w}_t. $$ -

    where \( \beta \) controls the averaging time of the second moment and is -typically taken to be about \( \beta=0.9 \), \( \eta_t \) is a learning rate -typically chosen to be \( 10^{-3} \), and \( \epsilon\sim 10^{-8} \) is a -small regularization constant to prevent divergences. Multiplication -and division by vectors is understood as an element-wise operation. It -is clear from this formula that the learning rate is reduced in -directions where the norm of the gradient is consistently large. This -greatly speeds up the convergence by allowing us to use a larger -learning rate for flat directions. -

    @@ -372,7 +385,7 @@ learning rate for flat directions.

  • 29
  • 30
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs021.html b/doc/pub/week40/html/._week40-bs021.html index 58a80c032..e9a0565b3 100644 --- a/doc/pub/week40/html/._week40-bs021.html +++ b/doc/pub/week40/html/._week40-bs021.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,59 +334,57 @@ MathJax.Hub.Config({

     

     

     

    -

    ADAM optimizer

    +

    Momentum parameter

    -

    A related algorithm is the ADAM optimizer. In -ADAM, we keep a running average of -both the first and second moment of the gradient and use this -information to adaptively change the learning rate for different -parameters. The method isefficient when working with large -problems involving lots data and/or parameters. It is a combination of the -gradient descent with momentum algorithm and the RMSprop algorithm -discussed above. +

    Notice that this equation is identical to previous one if we identify +the position of the particle, \( \mathbf{w} \), with the parameters +\( \boldsymbol{\theta} \). This allows us to identify the momentum +parameter and learning rate with the mass of the particle and the +viscous drag as:

    -

    In addition to keeping a running average of the first and -second moments of the gradient -(i.e. \( \mathbf{m}_t=\mathbb{E}[\mathbf{g}_t] \) and -\( \mathbf{s}_t=\mathbb{E}[\mathbf{g}^2_t] \), respectively), ADAM -performs an additional bias correction to account for the fact that we -are estimating the first two moments of the gradient using a running -average (denoted by the hats in the update rule below). The update -rule for ADAM is given by (where multiplication and division are once -again understood to be element-wise operations below) +$$ +\gamma= {m \over m +\mu \Delta t }, \qquad \eta = {(\Delta t)^2 \over m +\mu \Delta t}. +$$ + +

    Thus, as the name suggests, the momentum parameter is proportional to +the mass of the particle and effectively provides inertia. +Furthermore, in the large viscosity/small learning rate limit, our +memory time scales as \( (1-\gamma)^{-1} \approx m/(\mu \Delta t) \). +

    + +

    Why is momentum useful? SGD momentum helps the gradient descent +algorithm gain speed in directions with persistent but small gradients +even in the presence of stochasticity, while suppressing oscillations +in high-curvature directions. This becomes especially important in +situations where the landscape is shallow and flat in some directions +and narrow and steep in others. It has been argued that first-order +methods (with appropriate initial conditions) can perform comparable +to more expensive second order methods, especially in the context of +complex deep learning models. +

    + +

    These beneficial properties of momentum can sometimes become even more +pronounced by using a slight modification of the classical momentum +algorithm called Nesterov Accelerated Gradient (NAG). +

    + +

    In the NAG algorithm, rather than calculating the gradient at the +current parameters, \( \nabla_\theta E(\boldsymbol{\theta}_t) \), one +calculates the gradient at the expected value of the parameters given +our current momentum, \( \nabla_\theta E(\boldsymbol{\theta}_t +\gamma +\mathbf{v}_{t-1}) \). This yields the NAG update rule

    $$ \begin{align} -\mathbf{g}_t &= \nabla_\theta E(\boldsymbol{\theta}) -\tag{4}\\ -\mathbf{m}_t &= \beta_1 \mathbf{m}_{t-1} + (1-\beta_1) \mathbf{g}_t \nonumber \\ -\mathbf{s}_t &=\beta_2 \mathbf{s}_{t-1} +(1-\beta_2)\mathbf{g}_t^2 \nonumber \\ -\boldsymbol{\mathbf{m}}_t&={\mathbf{m}_t \over 1-\beta_1^t} \nonumber \\ -\boldsymbol{\mathbf{s}}_t &={\mathbf{s}_t \over1-\beta_2^t} \nonumber \\ -\boldsymbol{\theta}_{t+1}&=\boldsymbol{\theta}_t - \eta_t { \boldsymbol{\mathbf{m}}_t \over \sqrt{\boldsymbol{\mathbf{s}}_t} +\epsilon}, \nonumber \\ -\tag{5} +\mathbf{v}_{t}&=\gamma \mathbf{v}_{t-1}+\eta_{t}\nabla_\theta E(\boldsymbol{\theta}_t +\gamma \mathbf{v}_{t-1}) \nonumber \\ +\boldsymbol{\theta}_{t+1}&= \boldsymbol{\theta}_t -\mathbf{v}_{t}. +\tag{2} \end{align} $$ -

    where \( \beta_1 \) and \( \beta_2 \) set the memory lifetime of the first and -second moment and are typically taken to be \( 0.9 \) and \( 0.99 \) -respectively, and \( \eta \) and \( \epsilon \) are identical to RMSprop. -

    - -

    Like in RMSprop, the effective step size of a parameter depends on the -magnitude of its gradient squared. To understand this better, let us -rewrite this expression in terms of the variance -\( \boldsymbol{\sigma}_t^2 = \boldsymbol{\mathbf{s}}_t - -(\boldsymbol{\mathbf{m}}_t)^2 \). Consider a single parameter \( \theta_t \). The -update rule for this parameter is given by -

    - -$$ -\Delta \theta_{t+1}= -\eta_t { \boldsymbol{m}_t \over \sqrt{\sigma_t^2 + m_t^2 }+\epsilon}. -$$ - +

    One of the major advantages of NAG is that it allows for the use of a larger learning rate than GDM for the same choice of \( \gamma \).

    @@ -398,7 +411,7 @@ $$

  • 30
  • 31
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs022.html b/doc/pub/week40/html/._week40-bs022.html index 45d096092..ad90aa158 100644 --- a/doc/pub/week40/html/._week40-bs022.html +++ b/doc/pub/week40/html/._week40-bs022.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,11 +334,30 @@ MathJax.Hub.Config({

     

     

     

    -

    Algorithms and codes for Adagrad, RMSprop and Adam

    +

    Second moment of the gradient

    -

    The algorithms we have implemented are well described in the text by Goodfellow, Bengio and Courville, chapter 8.

    +

    In stochastic gradient descent, with and without momentum, we still +have to specify a schedule for tuning the learning rates \( \eta_t \) +as a function of time. As discussed in the context of Newton's +method, this presents a number of dilemmas. The learning rate is +limited by the steepest direction which can change depending on the +current position in the landscape. To circumvent this problem, ideally +our algorithm would keep track of curvature and take large steps in +shallow, flat directions and small steps in steep, narrow directions. +Second-order methods accomplish this by calculating or approximating +the Hessian and normalizing the learning rate by the +curvature. However, this is very computationally expensive for +extremely large models. Ideally, we would like to be able to +adaptively change the step size to match the landscape without paying +the steep computational price of calculating or approximating +Hessians. +

    -

    The codes which implement these algorithms are discussed after our presentation of automatic differentiation.

    +

    During the last decade a number of methods have been introduced that accomplish +this by tracking not only the gradient, but also the second moment of +the gradient. These methods include AdaGrad, AdaDelta, Root Mean Squared Propagation (RMS-Prop), and +ADAM. +

    @@ -350,7 +384,7 @@ MathJax.Hub.Config({

  • 31
  • 32
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs023.html b/doc/pub/week40/html/._week40-bs023.html index b11f3de28..53c12c676 100644 --- a/doc/pub/week40/html/._week40-bs023.html +++ b/doc/pub/week40/html/._week40-bs023.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,15 +334,33 @@ MathJax.Hub.Config({

     

     

     

    -

    Practical tips

    +

    RMS prop

    - -

    Geron's text, see chapter 11, has several interesting discussions.

    +

    In RMS prop, in addition to keeping a running average of the first +moment of the gradient, we also keep track of the second moment +denoted by \( \mathbf{s}_t=\mathbb{E}[\mathbf{g}_t^2] \). The update rule +for RMS prop is given by +

    + +$$ +\begin{align} +\mathbf{g}_t &= \nabla_\theta E(\boldsymbol{\theta}) +\tag{3}\\ +\mathbf{s}_t &=\beta \mathbf{s}_{t-1} +(1-\beta)\mathbf{g}_t^2 \nonumber \\ +\boldsymbol{\theta}_{t+1}&=&\boldsymbol{\theta}_t - \eta_t { \mathbf{g}_t \over \sqrt{\mathbf{s}_t +\epsilon}}, \nonumber +\end{align} +$$ + +

    where \( \beta \) controls the averaging time of the second moment and is +typically taken to be about \( \beta=0.9 \), \( \eta_t \) is a learning rate +typically chosen to be \( 10^{-3} \), and \( \epsilon\sim 10^{-8} \) is a +small regularization constant to prevent divergences. Multiplication +and division by vectors is understood as an element-wise operation. It +is clear from this formula that the learning rate is reduced in +directions where the norm of the gradient is consistently large. This +greatly speeds up the convergence by allowing us to use a larger +learning rate for flat directions. +

    @@ -354,7 +387,7 @@ MathJax.Hub.Config({

  • 32
  • 33
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs024.html b/doc/pub/week40/html/._week40-bs024.html index 6810e7037..6b2ca8049 100644 --- a/doc/pub/week40/html/._week40-bs024.html +++ b/doc/pub/week40/html/._week40-bs024.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,105 +334,59 @@ MathJax.Hub.Config({

     

     

     

    -

    Automatic differentiation

    +

    ADAM optimizer

    -

    Automatic differentiation (AD), -also called algorithmic -differentiation or computational differentiation,is a set of -techniques to numerically evaluate the derivative of a function -specified by a computer program. AD exploits the fact that every -computer program, no matter how complicated, executes a sequence of -elementary arithmetic operations (addition, subtraction, -multiplication, division, etc.) and elementary functions (exp, log, -sin, cos, etc.). By applying the chain rule repeatedly to these -operations, derivatives of arbitrary order can be computed -automatically, accurately to working precision, and using at most a -small constant factor more arithmetic operations than the original -program. +

    A related algorithm is the ADAM optimizer. In +ADAM, we keep a running average of +both the first and second moment of the gradient and use this +information to adaptively change the learning rate for different +parameters. The method isefficient when working with large +problems involving lots data and/or parameters. It is a combination of the +gradient descent with momentum algorithm and the RMSprop algorithm +discussed above.

    -

    Automatic differentiation is neither:

    - - -

    Symbolic differentiation can lead to inefficient code and faces the -difficulty of converting a computer program into a single expression, -while numerical differentiation can introduce round-off errors in the -discretization process and cancellation +

    In addition to keeping a running average of the first and +second moments of the gradient +(i.e. \( \mathbf{m}_t=\mathbb{E}[\mathbf{g}_t] \) and +\( \mathbf{s}_t=\mathbb{E}[\mathbf{g}^2_t] \), respectively), ADAM +performs an additional bias correction to account for the fact that we +are estimating the first two moments of the gradient using a running +average (denoted by the hats in the update rule below). The update +rule for ADAM is given by (where multiplication and division are once +again understood to be element-wise operations below)

    -

    Python has tools for so-called automatic differentiation. -Consider the following example +$$ +\begin{align} +\mathbf{g}_t &= \nabla_\theta E(\boldsymbol{\theta}) +\tag{4}\\ +\mathbf{m}_t &= \beta_1 \mathbf{m}_{t-1} + (1-\beta_1) \mathbf{g}_t \nonumber \\ +\mathbf{s}_t &=\beta_2 \mathbf{s}_{t-1} +(1-\beta_2)\mathbf{g}_t^2 \nonumber \\ +\boldsymbol{\mathbf{m}}_t&={\mathbf{m}_t \over 1-\beta_1^t} \nonumber \\ +\boldsymbol{\mathbf{s}}_t &={\mathbf{s}_t \over1-\beta_2^t} \nonumber \\ +\boldsymbol{\theta}_{t+1}&=\boldsymbol{\theta}_t - \eta_t { \boldsymbol{\mathbf{m}}_t \over \sqrt{\boldsymbol{\mathbf{s}}_t} +\epsilon}, \nonumber \\ +\tag{5} +\end{align} +$$ + +

    where \( \beta_1 \) and \( \beta_2 \) set the memory lifetime of the first and +second moment and are typically taken to be \( 0.9 \) and \( 0.99 \) +respectively, and \( \eta \) and \( \epsilon \) are identical to RMSprop.

    + +

    Like in RMSprop, the effective step size of a parameter depends on the +magnitude of its gradient squared. To understand this better, let us +rewrite this expression in terms of the variance +\( \boldsymbol{\sigma}_t^2 = \boldsymbol{\mathbf{s}}_t - +(\boldsymbol{\mathbf{m}}_t)^2 \). Consider a single parameter \( \theta_t \). The +update rule for this parameter is given by +

    + $$ -f(x) = \sin\left(2\pi x + x^2\right) +\Delta \theta_{t+1}= -\eta_t { \boldsymbol{m}_t \over \sqrt{\sigma_t^2 + m_t^2 }+\epsilon}. $$ -

    which has the following derivative

    -$$ -f'(x) = \cos\left(2\pi x + x^2\right)\left(2\pi + 2x\right) -$$ - -

    Using autograd we have

    - - - -
    -
    -
    -
    -
    -
    import autograd.numpy as np
    -
    -# To do elementwise differentiation:
    -from autograd import elementwise_grad as egrad 
    -
    -# To plot:
    -import matplotlib.pyplot as plt 
    -
    -
    -def f(x):
    -    return np.sin(2*np.pi*x + x**2)
    -
    -def f_grad_analytic(x):
    -    return np.cos(2*np.pi*x + x**2)*(2*np.pi + 2*x)
    -
    -# Do the comparison:
    -x = np.linspace(0,1,1000)
    -
    -f_grad = egrad(f)
    -
    -computed = f_grad(x)
    -analytic = f_grad_analytic(x)
    -
    -plt.title('Derivative computed from Autograd compared with the analytical derivative')
    -plt.plot(x,computed,label='autograd')
    -plt.plot(x,analytic,label='analytic')
    -
    -plt.xlabel('x')
    -plt.ylabel('y')
    -plt.legend()
    -
    -plt.show()
    -
    -print("The max absolute difference is: %g"%(np.max(np.abs(computed - analytic))))
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -

    @@ -444,7 +413,7 @@ plt.show()

  • 33
  • 34
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs025.html b/doc/pub/week40/html/._week40-bs025.html index 1bfa65577..cfa83373c 100644 --- a/doc/pub/week40/html/._week40-bs025.html +++ b/doc/pub/week40/html/._week40-bs025.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -318,55 +333,12 @@ MathJax.Hub.Config({

     

     

     

    - -

    Using autograd

    + +

    Algorithms and codes for Adagrad, RMSprop and Adam

    -

    Here we -experiment with what kind of functions Autograd is capable -of finding the gradient of. The following Python functions are just -meant to illustrate what Autograd can do, but please feel free to -experiment with other, possibly more complicated, functions as well. -

    - - - -
    -
    -
    -
    -
    -
    import autograd.numpy as np
    -from autograd import grad
    -
    -def f1(x):
    -    return x**3 + 1
    -
    -f1_grad = grad(f1)
    -
    -# Remember to send in float as argument to the computed gradient from Autograd!
    -a = 1.0
    -
    -# See the evaluated gradient at a using autograd:
    -print("The gradient of f1 evaluated at a = %g using autograd is: %g"%(a,f1_grad(a)))
    -
    -# Compare with the analytical derivative, that is f1'(x) = 3*x**2 
    -grad_analytical = 3*a**2
    -print("The gradient of f1 evaluated at a = %g by finding the analytic expression is: %g"%(a,grad_analytical))
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    +

    The algorithms we have implemented are well described in the text by Goodfellow, Bengio and Courville, chapter 8.

    +

    The codes which implement these algorithms are discussed after our presentation of automatic differentiation.

    @@ -393,7 +365,7 @@ grad_analytical = 34

  • 35
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs026.html b/doc/pub/week40/html/._week40-bs026.html index f62f1327c..8bade0c65 100644 --- a/doc/pub/week40/html/._week40-bs026.html +++ b/doc/pub/week40/html/._week40-bs026.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,69 +334,15 @@ MathJax.Hub.Config({

     

     

     

    -

    Autograd with more complicated functions

    +

    Practical tips

    -

    To differentiate with respect to two (or more) arguments of a Python -function, Autograd need to know at which variable the function if -being differentiated with respect to. -

    - - - -
    -
    -
    -
    -
    -
    import autograd.numpy as np
    -from autograd import grad
    -def f2(x1,x2):
    -    return 3*x1**3 + x2*(x1 - 5) + 1
    -
    -# By sending the argument 0, Autograd will compute the derivative w.r.t the first variable, in this case x1
    -f2_grad_x1 = grad(f2,0)
    -
    -# ... and differentiate w.r.t x2 by sending 1 as an additional arugment to grad
    -f2_grad_x2 = grad(f2,1)
    -
    -x1 = 1.0
    -x2 = 3.0 
    -
    -print("Evaluating at x1 = %g, x2 = %g"%(x1,x2))
    -print("-"*30)
    -
    -# Compare with the analytical derivatives:
    -
    -# Derivative of f2 w.r.t x1 is: 9*x1**2 + x2:
    -f2_grad_x1_analytical = 9*x1**2 + x2
    -
    -# Derivative of f2 w.r.t x2 is: x1 - 5:
    -f2_grad_x2_analytical = x1 - 5
    -
    -# See the evaluated derivations:
    -print("The derivative of f2 w.r.t x1: %g"%( f2_grad_x1(x1,x2) ))
    -print("The analytical derivative of f2 w.r.t x1: %g"%( f2_grad_x1(x1,x2) ))
    -
    -print()
    -
    -print("The derivative of f2 w.r.t x2: %g"%( f2_grad_x2(x1,x2) ))
    -print("The analytical derivative of f2 w.r.t x2: %g"%( f2_grad_x2(x1,x2) ))
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - -

    Note that the grad function will not produce the true gradient of the function. The true gradient of a function with two or more variables will produce a vector, where each element is the function differentiated w.r.t a variable.

    + +

    Geron's text, see chapter 11, has several interesting discussions.

    @@ -408,7 +369,7 @@ f2_grad_x2_analytical = x1 35

  • 36
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs027.html b/doc/pub/week40/html/._week40-bs027.html index 641ced72a..232f6d1aa 100644 --- a/doc/pub/week40/html/._week40-bs027.html +++ b/doc/pub/week40/html/._week40-bs027.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,7 +334,48 @@ MathJax.Hub.Config({

     

     

     

    -

    More complicated functions using the elements of their arguments directly

    +

    Automatic differentiation

    + +

    Automatic differentiation (AD), +also called algorithmic +differentiation or computational differentiation,is a set of +techniques to numerically evaluate the derivative of a function +specified by a computer program. AD exploits the fact that every +computer program, no matter how complicated, executes a sequence of +elementary arithmetic operations (addition, subtraction, +multiplication, division, etc.) and elementary functions (exp, log, +sin, cos, etc.). By applying the chain rule repeatedly to these +operations, derivatives of arbitrary order can be computed +automatically, accurately to working precision, and using at most a +small constant factor more arithmetic operations than the original +program. +

    + +

    Automatic differentiation is neither:

    + + +

    Symbolic differentiation can lead to inefficient code and faces the +difficulty of converting a computer program into a single expression, +while numerical differentiation can introduce round-off errors in the +discretization process and cancellation +

    + +

    Python has tools for so-called automatic differentiation. +Consider the following example +

    +$$ +f(x) = \sin\left(2\pi x + x^2\right) +$$ + +

    which has the following derivative

    +$$ +f'(x) = \cos\left(2\pi x + x^2\right)\left(2\pi + 2x\right) +$$ + +

    Using autograd we have

    @@ -329,22 +385,39 @@ MathJax.Hub.Config({
    import autograd.numpy as np
    -from autograd import grad
    -def f3(x): # Assumes x is an array of length 5 or higher
    -    return 2*x[0] + 3*x[1] + 5*x[2] + 7*x[3] + 11*x[4]**2
     
    -f3_grad = grad(f3)
    +# To do elementwise differentiation:
    +from autograd import elementwise_grad as egrad 
     
    -x = np.linspace(0,4,5)
    +# To plot:
    +import matplotlib.pyplot as plt 
     
    -# Print the computed gradient:
    -print("The computed gradient of f3 is: ", f3_grad(x))
     
    -# The analytical gradient is: (2, 3, 5, 7, 22*x[4])
    -f3_grad_analytical = np.array([2, 3, 5, 7, 22*x[4]])
    +def f(x):
    +    return np.sin(2*np.pi*x + x**2)
     
    -# Print the analytical gradient:
    -print("The analytical gradient of f3 is: ", f3_grad_analytical)
    +def f_grad_analytic(x):
    +    return np.cos(2*np.pi*x + x**2)*(2*np.pi + 2*x)
    +
    +# Do the comparison:
    +x = np.linspace(0,1,1000)
    +
    +f_grad = egrad(f)
    +
    +computed = f_grad(x)
    +analytic = f_grad_analytic(x)
    +
    +plt.title('Derivative computed from Autograd compared with the analytical derivative')
    +plt.plot(x,computed,label='autograd')
    +plt.plot(x,analytic,label='analytic')
    +
    +plt.xlabel('x')
    +plt.ylabel('y')
    +plt.legend()
    +
    +plt.show()
    +
    +print("The max absolute difference is: %g"%(np.max(np.abs(computed - analytic))))
     
    @@ -360,13 +433,6 @@ f3_grad_analytical = np36
  • 37
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs028.html b/doc/pub/week40/html/._week40-bs028.html index ce18f2ae7..aea64abe5 100644 --- a/doc/pub/week40/html/._week40-bs028.html +++ b/doc/pub/week40/html/._week40-bs028.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,7 +334,14 @@ MathJax.Hub.Config({

     

     

     

    -

    Functions using mathematical functions from Numpy

    +

    Using autograd

    + +

    Here we +experiment with what kind of functions Autograd is capable +of finding the gradient of. The following Python functions are just +meant to illustrate what Autograd can do, but please feel free to +experiment with other, possibly more complicated, functions as well. +

    @@ -330,21 +352,21 @@ MathJax.Hub.Config({
    import autograd.numpy as np
     from autograd import grad
    -def f4(x):
    -    return np.sqrt(1+x**2) + np.exp(x) + np.sin(2*np.pi*x)
     
    -f4_grad = grad(f4)
    +def f1(x):
    +    return x**3 + 1
     
    -x = 2.7
    +f1_grad = grad(f1)
     
    -# Print the computed derivative:
    -print("The computed derivative of f4 at x = %g is: %g"%(x,f4_grad(x)))
    +# Remember to send in float as argument to the computed gradient from Autograd!
    +a = 1.0
     
    -# The analytical derivative is: x/sqrt(1 + x**2) + exp(x) + cos(2*pi*x)*2*pi
    -f4_grad_analytical = x/np.sqrt(1 + x**2) + np.exp(x) + np.cos(2*np.pi*x)*2*np.pi
    +# See the evaluated gradient at a using autograd:
    +print("The gradient of f1 evaluated at a = %g using autograd is: %g"%(a,f1_grad(a)))
     
    -# Print the analytical gradient:
    -print("The analytical gradient of f4 at x = %g is: %g"%(x,f4_grad_analytical))
    +# Compare with the analytical derivative, that is f1'(x) = 3*x**2 
    +grad_analytical = 3*a**2
    +print("The gradient of f1 evaluated at a = %g by finding the analytic expression is: %g"%(a,grad_analytical))
     
    @@ -386,7 +408,7 @@ f4_grad_analytical = x37
  • 38
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs029.html b/doc/pub/week40/html/._week40-bs029.html index 0bf2d94f8..2ae8d8f76 100644 --- a/doc/pub/week40/html/._week40-bs029.html +++ b/doc/pub/week40/html/._week40-bs029.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,7 +334,12 @@ MathJax.Hub.Config({

     

     

     

    -

    More autograd

    +

    Autograd with more complicated functions

    + +

    To differentiate with respect to two (or more) arguments of a Python +function, Autograd need to know at which variable the function if +being differentiated with respect to. +

    @@ -330,18 +350,37 @@ MathJax.Hub.Config({
    import autograd.numpy as np
     from autograd import grad
    -def f5(x):
    -    if x >= 0:
    -        return x**2
    -    else:
    -        return -3*x + 1
    +def f2(x1,x2):
    +    return 3*x1**3 + x2*(x1 - 5) + 1
     
    -f5_grad = grad(f5)
    +# By sending the argument 0, Autograd will compute the derivative w.r.t the first variable, in this case x1
    +f2_grad_x1 = grad(f2,0)
     
    -x = 2.7
    +# ... and differentiate w.r.t x2 by sending 1 as an additional arugment to grad
    +f2_grad_x2 = grad(f2,1)
     
    -# Print the computed derivative:
    -print("The computed derivative of f5 at x = %g is: %g"%(x,f5_grad(x)))
    +x1 = 1.0
    +x2 = 3.0 
    +
    +print("Evaluating at x1 = %g, x2 = %g"%(x1,x2))
    +print("-"*30)
    +
    +# Compare with the analytical derivatives:
    +
    +# Derivative of f2 w.r.t x1 is: 9*x1**2 + x2:
    +f2_grad_x1_analytical = 9*x1**2 + x2
    +
    +# Derivative of f2 w.r.t x2 is: x1 - 5:
    +f2_grad_x2_analytical = x1 - 5
    +
    +# See the evaluated derivations:
    +print("The derivative of f2 w.r.t x1: %g"%( f2_grad_x1(x1,x2) ))
    +print("The analytical derivative of f2 w.r.t x1: %g"%( f2_grad_x1(x1,x2) ))
    +
    +print()
    +
    +print("The derivative of f2 w.r.t x2: %g"%( f2_grad_x2(x1,x2) ))
    +print("The analytical derivative of f2 w.r.t x2: %g"%( f2_grad_x2(x1,x2) ))
     
    @@ -357,6 +396,7 @@ x = 2.7 +

    Note that the grad function will not produce the true gradient of the function. The true gradient of a function with two or more variables will produce a vector, where each element is the function differentiated w.r.t a variable.

    @@ -383,7 +423,7 @@ x = 2.7

  • 38
  • 39
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs030.html b/doc/pub/week40/html/._week40-bs030.html index 29767aa57..e92995562 100644 --- a/doc/pub/week40/html/._week40-bs030.html +++ b/doc/pub/week40/html/._week40-bs030.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,7 +334,7 @@ MathJax.Hub.Config({

     

     

     

    -

    And with loops

    +

    More complicated functions using the elements of their arguments directly

    @@ -330,59 +345,21 @@ MathJax.Hub.Config({
    import autograd.numpy as np
     from autograd import grad
    -def f6_for(x):
    -    val = 0
    -    for i in range(10):
    -        val = val + x**i
    -    return val
    +def f3(x): # Assumes x is an array of length 5 or higher
    +    return 2*x[0] + 3*x[1] + 5*x[2] + 7*x[3] + 11*x[4]**2
     
    -def f6_while(x):
    -    val = 0
    -    i = 0
    -    while i < 10:
    -        val = val + x**i
    -        i = i + 1
    -    return val
    +f3_grad = grad(f3)
     
    -f6_for_grad = grad(f6_for)
    -f6_while_grad = grad(f6_while)
    +x = np.linspace(0,4,5)
     
    -x = 0.5
    +# Print the computed gradient:
    +print("The computed gradient of f3 is: ", f3_grad(x))
     
    -# Print the computed derivaties of f6_for and f6_while
    -print("The computed derivative of f6_for at x = %g is: %g"%(x,f6_for_grad(x)))
    -print("The computed derivative of f6_while at x = %g is: %g"%(x,f6_while_grad(x)))
    -
    -
    - - - -
    -
    -
    -
    -
    -
    -
    -
    - - - - -
    -
    -
    -
    -
    -
    import autograd.numpy as np
    -from autograd import grad
    -# Both of the functions are implementation of the sum: sum(x**i) for i = 0, ..., 9
    -# The analytical derivative is: sum(i*x**(i-1)) 
    -f6_grad_analytical = 0
    -for i in range(10):
    -    f6_grad_analytical += i*x**(i-1)
    -
    -print("The analytical derivative of f6 at x = %g is: %g"%(x,f6_grad_analytical))
    +# The analytical gradient is: (2, 3, 5, 7, 22*x[4])
    +f3_grad_analytical = np.array([2, 3, 5, 7, 22*x[4]])
    +
    +# Print the analytical gradient:
    +print("The analytical gradient of f3 is: ", f3_grad_analytical)
     
    @@ -398,6 +375,13 @@ f6_grad_analytical = = 39
  • 40
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs031.html b/doc/pub/week40/html/._week40-bs031.html index c0d5990f8..0c076d805 100644 --- a/doc/pub/week40/html/._week40-bs031.html +++ b/doc/pub/week40/html/._week40-bs031.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -318,8 +333,9 @@ MathJax.Hub.Config({

     

     

     

    - -

    Using recursion

    + +

    Functions using mathematical functions from Numpy

    +
    @@ -329,31 +345,21 @@ MathJax.Hub.Config({
    import autograd.numpy as np
     from autograd import grad
    +def f4(x):
    +    return np.sqrt(1+x**2) + np.exp(x) + np.sin(2*np.pi*x)
     
    -def f7(n): # Assume that n is an integer
    -    if n == 1 or n == 0:
    -        return 1
    -    else:
    -        return n*f7(n-1)
    +f4_grad = grad(f4)
     
    -f7_grad = grad(f7)
    +x = 2.7
     
    -n = 2.0
    +# Print the computed derivative:
    +print("The computed derivative of f4 at x = %g is: %g"%(x,f4_grad(x)))
     
    -print("The computed derivative of f7 at n = %d is: %g"%(n,f7_grad(n)))
    +# The analytical derivative is: x/sqrt(1 + x**2) + exp(x) + cos(2*pi*x)*2*pi
    +f4_grad_analytical = x/np.sqrt(1 + x**2) + np.exp(x) + np.cos(2*np.pi*x)*2*np.pi
     
    -# The function f7 is an implementation of the factorial of n.
    -# By using the product rule, one can find that the derivative is:
    -
    -f7_grad_analytical = 0
    -for i in range(int(n)-1):
    -    tmp = 1
    -    for k in range(int(n)-1):
    -        if k != i:
    -            tmp *= (n - k)
    -    f7_grad_analytical += tmp
    -
    -print("The analytical derivative of f7 at n = %d is: %g"%(n,f7_grad_analytical))
    +# Print the analytical gradient:
    +print("The analytical gradient of f4 at x = %g is: %g"%(x,f4_grad_analytical))
     
    @@ -369,7 +375,6 @@ f7_grad_analytical = = 40
  • 41
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs032.html b/doc/pub/week40/html/._week40-bs032.html index 620551e5f..f5fb34795 100644 --- a/doc/pub/week40/html/._week40-bs032.html +++ b/doc/pub/week40/html/._week40-bs032.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,10 +334,8 @@ MathJax.Hub.Config({

     

     

     

    -

    Unsupported functions

    -

    Autograd supports many features. However, there are some functions that is not supported (yet) by Autograd.

    +

    More autograd

    -

    Assigning a value to the variable being differentiated with respect to

    @@ -332,15 +345,18 @@ MathJax.Hub.Config({
    import autograd.numpy as np
     from autograd import grad
    -def f8(x): # Assume x is an array
    -    x[2] = 3
    -    return x*2
    +def f5(x):
    +    if x >= 0:
    +        return x**2
    +    else:
    +        return -3*x + 1
     
    -f8_grad = grad(f8)
    +f5_grad = grad(f5)
     
    -x = 8.4
    +x = 2.7
     
    -print("The derivative of f8 is:",f8_grad(x))
    +# Print the computed derivative:
    +print("The computed derivative of f5 at x = %g is: %g"%(x,f5_grad(x)))
     
    @@ -356,7 +372,6 @@ x = 8.4
    -

    Here, Autograd tells us that an 'ArrayBox' does not support item assignment. The item assignment is done when the program tries to assign x[2] to the value 3. However, Autograd has implemented the computation of the derivative such that this assignment is not possible.

    @@ -383,7 +398,7 @@ x = 8.4

  • 41
  • 42
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs033.html b/doc/pub/week40/html/._week40-bs033.html index 2022e3031..9e354f45d 100644 --- a/doc/pub/week40/html/._week40-bs033.html +++ b/doc/pub/week40/html/._week40-bs033.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,7 +334,8 @@ MathJax.Hub.Config({

     

     

     

    -

    The syntax a.dot(b) when finding the dot product

    +

    And with loops

    +
    @@ -329,15 +345,28 @@ MathJax.Hub.Config({
    import autograd.numpy as np
     from autograd import grad
    -def f9(a): # Assume a is an array with 2 elements
    -    b = np.array([1.0,2.0])
    -    return a.dot(b)
    +def f6_for(x):
    +    val = 0
    +    for i in range(10):
    +        val = val + x**i
    +    return val
     
    -f9_grad = grad(f9)
    +def f6_while(x):
    +    val = 0
    +    i = 0
    +    while i < 10:
    +        val = val + x**i
    +        i = i + 1
    +    return val
     
    -x = np.array([1.0,0.0])
    +f6_for_grad = grad(f6_for)
    +f6_while_grad = grad(f6_while)
     
    -print("The derivative of f9 is:",f9_grad(x))
    +x = 0.5
    +
    +# Print the computed derivaties of f6_for and f6_while
    +print("The computed derivative of f6_for at x = %g is: %g"%(x,f6_for_grad(x)))
    +print("The computed derivative of f6_while at x = %g is: %g"%(x,f6_while_grad(x)))
     
    @@ -353,11 +382,6 @@ x = np.a
    -

    Here we are told that the 'dot' function does not belong to Autograd's -version of a Numpy array. To overcome this, an alternative syntax -which also computed the dot product can be used: -

    -
    @@ -367,18 +391,13 @@ which also computed the dot product can be used:
    import autograd.numpy as np
     from autograd import grad
    -def f9_alternative(x): # Assume a is an array with 2 elements
    -    b = np.array([1.0,2.0])
    -    return np.dot(x,b) # The same as x_1*b_1 + x_2*b_2
    +# Both of the functions are implementation of the sum: sum(x**i) for i = 0, ..., 9
    +# The analytical derivative is: sum(i*x**(i-1)) 
    +f6_grad_analytical = 0
    +for i in range(10):
    +    f6_grad_analytical += i*x**(i-1)
     
    -f9_alternative_grad = grad(f9_alternative)
    -
    -x = np.array([3.0,0.0])
    -
    -print("The gradient of f9 is:",f9_alternative_grad(x))
    -
    -# The analytical gradient of the dot product of vectors x and b with two elements (x_1,x_2) and (b_1, b_2) respectively
    -# w.r.t x is (b_1, b_2).
    +print("The analytical derivative of f6 at x = %g is: %g"%(x,f6_grad_analytical))
     
    @@ -420,7 +439,7 @@ x = np.a
  • 42
  • 43
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs034.html b/doc/pub/week40/html/._week40-bs034.html index 9babf2712..5a08c85aa 100644 --- a/doc/pub/week40/html/._week40-bs034.html +++ b/doc/pub/week40/html/._week40-bs034.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,8 +334,7 @@ MathJax.Hub.Config({

     

     

     

    - -

    The documentation recommends to avoid inplace operations such as

    +

    Using recursion

    @@ -328,10 +342,33 @@ MathJax.Hub.Config({
    -
    a += b
    -a -= b
    -a*= b
    -a /=b
    +  
    import autograd.numpy as np
    +from autograd import grad
    +
    +def f7(n): # Assume that n is an integer
    +    if n == 1 or n == 0:
    +        return 1
    +    else:
    +        return n*f7(n-1)
    +
    +f7_grad = grad(f7)
    +
    +n = 2.0
    +
    +print("The computed derivative of f7 at n = %d is: %g"%(n,f7_grad(n)))
    +
    +# The function f7 is an implementation of the factorial of n.
    +# By using the product rule, one can find that the derivative is:
    +
    +f7_grad_analytical = 0
    +for i in range(int(n)-1):
    +    tmp = 1
    +    for k in range(int(n)-1):
    +        if k != i:
    +            tmp *= (n - k)
    +    f7_grad_analytical += tmp
    +
    +print("The analytical derivative of f7 at n = %d is: %g"%(n,f7_grad_analytical))
     
    @@ -347,6 +384,7 @@ a /=b
    +

    Note that if n is equal to zero or one, Autograd will give an error message. This message appears when the output is independent on input.

    @@ -373,7 +411,7 @@ a /=b

  • 43
  • 44
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs035.html b/doc/pub/week40/html/._week40-bs035.html index 5d926ac4d..1fb7bfc00 100644 --- a/doc/pub/week40/html/._week40-bs035.html +++ b/doc/pub/week40/html/._week40-bs035.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,13 +334,10 @@ MathJax.Hub.Config({

     

     

     

    -

    Using Autograd with OLS

    - -

    We conclude the part on optmization by showing how we can make codes -for linear regression and logistic regression using autograd. The -first example shows results with ordinary leats squares. -

    +

    Unsupported functions

    +

    Autograd supports many features. However, there are some functions that is not supported (yet) by Autograd.

    +

    Assigning a value to the variable being differentiated with respect to

    @@ -333,55 +345,17 @@ first example shows results with ordinary leats squares.
    -
    # Using Autograd to calculate gradients for OLS
    -from random import random, seed
    -import numpy as np
    -import autograd.numpy as np
    -import matplotlib.pyplot as plt
    +  
    import autograd.numpy as np
     from autograd import grad
    +def f8(x): # Assume x is an array
    +    x[2] = 3
    +    return x*2
     
    -def CostOLS(beta):
    -    return (1.0/n)*np.sum((y-X @ beta)**2)
    +f8_grad = grad(f8)
     
    -n = 100
    -x = 2*np.random.rand(n,1)
    -y = 4+3*x+np.random.randn(n,1)
    +x = 8.4
     
    -X = np.c_[np.ones((n,1)), x]
    -XT_X = X.T @ X
    -theta_linreg = np.linalg.pinv(XT_X) @ (X.T @ y)
    -print("Own inversion")
    -print(theta_linreg)
    -# Hessian matrix
    -H = (2.0/n)* XT_X
    -EigValues, EigVectors = np.linalg.eig(H)
    -print(f"Eigenvalues of Hessian Matrix:{EigValues}")
    -
    -theta = np.random.randn(2,1)
    -eta = 1.0/np.max(EigValues)
    -Niterations = 1000
    -# define the gradient
    -training_gradient = grad(CostOLS)
    -
    -for iter in range(Niterations):
    -    gradients = training_gradient(theta)
    -    theta -= eta*gradients
    -print("theta from own gd")
    -print(theta)
    -
    -xnew = np.array([[0],[2]])
    -Xnew = np.c_[np.ones((2,1)), xnew]
    -ypredict = Xnew.dot(theta)
    -ypredict2 = Xnew.dot(theta_linreg)
    -
    -plt.plot(xnew, ypredict, "r-")
    -plt.plot(xnew, ypredict2, "b-")
    -plt.plot(x, y ,'ro')
    -plt.axis([0,2.0,0, 15.0])
    -plt.xlabel(r'$x$')
    -plt.ylabel(r'$y$')
    -plt.title(r'Random numbers ')
    -plt.show()
    +print("The derivative of f8 is:",f8_grad(x))
     
    @@ -397,6 +371,7 @@ plt.show()
    +

    Here, Autograd tells us that an 'ArrayBox' does not support item assignment. The item assignment is done when the program tries to assign x[2] to the value 3. However, Autograd has implemented the computation of the derivative such that this assignment is not possible.

    @@ -423,7 +398,7 @@ plt.show()

  • 44
  • 45
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs036.html b/doc/pub/week40/html/._week40-bs036.html index fa409740a..55f48462b 100644 --- a/doc/pub/week40/html/._week40-bs036.html +++ b/doc/pub/week40/html/._week40-bs036.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,7 +334,7 @@ MathJax.Hub.Config({

     

     

     

    -

    Same code but now with momentum gradient descent

    +

    The syntax a.dot(b) when finding the dot product

    @@ -327,59 +342,58 @@ MathJax.Hub.Config({
    -
    # Using Autograd to calculate gradients for OLS
    -from random import random, seed
    -import numpy as np
    -import autograd.numpy as np
    -import matplotlib.pyplot as plt
    +  
    import autograd.numpy as np
     from autograd import grad
    +def f9(a): # Assume a is an array with 2 elements
    +    b = np.array([1.0,2.0])
    +    return a.dot(b)
     
    -def CostOLS(beta):
    -    return (1.0/n)*np.sum((y-X @ beta)**2)
    +f9_grad = grad(f9)
     
    -n = 100
    -x = 2*np.random.rand(n,1)
    -y = 4+3*x#+np.random.randn(n,1)
    +x = np.array([1.0,0.0])
     
    -X = np.c_[np.ones((n,1)), x]
    -XT_X = X.T @ X
    -theta_linreg = np.linalg.pinv(XT_X) @ (X.T @ y)
    -print("Own inversion")
    -print(theta_linreg)
    -# Hessian matrix
    -H = (2.0/n)* XT_X
    -EigValues, EigVectors = np.linalg.eig(H)
    -print(f"Eigenvalues of Hessian Matrix:{EigValues}")
    +print("The derivative of f9 is:",f9_grad(x))
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    + -theta = np.random.randn(2,1) -eta = 1.0/np.max(EigValues) -Niterations = 30 +

    Here we are told that the 'dot' function does not belong to Autograd's +version of a Numpy array. To overcome this, an alternative syntax +which also computed the dot product can be used: +

    -# define the gradient -training_gradient = grad(CostOLS) -for iter in range(Niterations): - gradients = training_gradient(theta) - theta -= eta*gradients - print(iter,gradients[0],gradients[1]) -print("theta from own gd") -print(theta) + +
    +
    +
    +
    +
    +
    import autograd.numpy as np
    +from autograd import grad
    +def f9_alternative(x): # Assume a is an array with 2 elements
    +    b = np.array([1.0,2.0])
    +    return np.dot(x,b) # The same as x_1*b_1 + x_2*b_2
     
    -# Now improve with momentum gradient descent
    -change = 0.0
    -delta_momentum = 0.3
    -for iter in range(Niterations):
    -    # calculate gradient
    -    gradients = training_gradient(theta)
    -    # calculate update
    -    new_change = eta*gradients+delta_momentum*change
    -    # take a step
    -    theta -= new_change
    -    # save the change
    -    change = new_change
    -    print(iter,gradients[0],gradients[1])
    -print("theta from own gd wth momentum")
    -print(theta)
    +f9_alternative_grad = grad(f9_alternative)
    +
    +x = np.array([3.0,0.0])
    +
    +print("The gradient of f9 is:",f9_alternative_grad(x))
    +
    +# The analytical gradient of the dot product of vectors x and b with two elements (x_1,x_2) and (b_1, b_2) respectively
    +# w.r.t x is (b_1, b_2).
     
    @@ -421,7 +435,7 @@ delta_momentum = 45
  • 46
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs037.html b/doc/pub/week40/html/._week40-bs037.html index f649a9615..ae2fac72b 100644 --- a/doc/pub/week40/html/._week40-bs037.html +++ b/doc/pub/week40/html/._week40-bs037.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,8 +334,8 @@ MathJax.Hub.Config({

     

     

     

    -

    But noen of these can compete with Newton's method

    - + +

    The documentation recommends to avoid inplace operations such as

    @@ -328,44 +343,10 @@ MathJax.Hub.Config({
    -
    # Using Newton's method
    -from random import random, seed
    -import numpy as np
    -import autograd.numpy as np
    -import matplotlib.pyplot as plt
    -from autograd import grad
    -
    -def CostOLS(beta):
    -    return (1.0/n)*np.sum((y-X @ beta)**2)
    -
    -n = 100
    -x = 2*np.random.rand(n,1)
    -y = 4+3*x+np.random.randn(n,1)
    -
    -X = np.c_[np.ones((n,1)), x]
    -XT_X = X.T @ X
    -beta_linreg = np.linalg.pinv(XT_X) @ (X.T @ y)
    -print("Own inversion")
    -print(beta_linreg)
    -# Hessian matrix
    -H = (2.0/n)* XT_X
    -# Note that here the Hessian does not depend on the parameters beta
    -invH = np.linalg.pinv(H)
    -EigValues, EigVectors = np.linalg.eig(H)
    -print(f"Eigenvalues of Hessian Matrix:{EigValues}")
    -
    -beta = np.random.randn(2,1)
    -Niterations = 5
    -
    -# define the gradient
    -training_gradient = grad(CostOLS)
    -
    -for iter in range(Niterations):
    -    gradients = training_gradient(beta)
    -    beta -= invH @ gradients
    -    print(iter,gradients[0],gradients[1])
    -print("beta from own Newton code")
    -print(beta)
    +  
    a += b
    +a -= b
    +a*= b
    +a /=b
     
    @@ -407,7 +388,7 @@ training_gradient = grad(CostOLS)
  • 46
  • 47
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs038.html b/doc/pub/week40/html/._week40-bs038.html index 7db2d6054..b915f8322 100644 --- a/doc/pub/week40/html/._week40-bs038.html +++ b/doc/pub/week40/html/._week40-bs038.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,8 +334,12 @@ MathJax.Hub.Config({

     

     

     

    -

    Including Stochastic Gradient Descent with Autograd

    -

    In this code we include the stochastic gradient descent approach discussed above. Note here that we specify which argument we are taking the derivative with respect to when using autograd.

    +

    Using Autograd with OLS

    + +

    We conclude the part on optmization by showing how we can make codes +for linear regression and logistic regression using autograd. The +first example shows results with ordinary leats squares. +

    @@ -329,17 +348,15 @@ MathJax.Hub.Config({
    -
    # Using Autograd to calculate gradients using SGD
    -# OLS example
    +  
    # Using Autograd to calculate gradients for OLS
     from random import random, seed
     import numpy as np
     import autograd.numpy as np
     import matplotlib.pyplot as plt
     from autograd import grad
     
    -# Note change from previous example
    -def CostOLS(y,X,theta):
    -    return np.sum((y-X @ theta)**2)
    +def CostOLS(beta):
    +    return (1.0/n)*np.sum((y-X @ beta)**2)
     
     n = 100
     x = 2*np.random.rand(n,1)
    @@ -358,12 +375,11 @@ EigValues, EigVectors = np= np.random.randn(2,1)
     eta = 1.0/np.max(EigValues)
     Niterations = 1000
    -
    -# Note that we request the derivative wrt third argument (theta, 2 here)
    -training_gradient = grad(CostOLS,2)
    +# define the gradient
    +training_gradient = grad(CostOLS)
     
     for iter in range(Niterations):
    -    gradients = (1.0/n)*training_gradient(y, X, theta)
    +    gradients = training_gradient(theta)
         theta -= eta*gradients
     print("theta from own gd")
     print(theta)
    @@ -381,27 +397,6 @@ plt.xlabel(r
     plt.ylabel(r'$y$')
     plt.title(r'Random numbers ')
     plt.show()
    -
    -n_epochs = 50
    -M = 5   #size of each minibatch
    -m = int(n/M) #number of minibatches
    -t0, t1 = 5, 50
    -def learning_schedule(t):
    -    return t0/(t+t1)
    -
    -theta = np.random.randn(2,1)
    -
    -for epoch in range(n_epochs):
    -# Can you figure out a better way of setting up the contributions to each batch?
    -    for i in range(m):
    -        random_index = M*np.random.randint(m)
    -        xi = X[random_index:random_index+M]
    -        yi = y[random_index:random_index+M]
    -        gradients = (1.0/M)*training_gradient(yi, xi, theta)
    -        eta = learning_schedule(epoch*m+i)
    -        theta = theta - eta*gradients
    -print("theta from own sdg")
    -print(theta)
     
    @@ -443,7 +438,7 @@ theta = np.47
  • 48
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs039.html b/doc/pub/week40/html/._week40-bs039.html index f2019bd8e..71e644bcf 100644 --- a/doc/pub/week40/html/._week40-bs039.html +++ b/doc/pub/week40/html/._week40-bs039.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -327,21 +342,19 @@ MathJax.Hub.Config({
    -
    # Using Autograd to calculate gradients using SGD
    -# OLS example
    +  
    # Using Autograd to calculate gradients for OLS
     from random import random, seed
     import numpy as np
     import autograd.numpy as np
     import matplotlib.pyplot as plt
     from autograd import grad
     
    -# Note change from previous example
    -def CostOLS(y,X,theta):
    -    return np.sum((y-X @ theta)**2)
    +def CostOLS(beta):
    +    return (1.0/n)*np.sum((y-X @ beta)**2)
     
     n = 100
     x = 2*np.random.rand(n,1)
    -y = 4+3*x+np.random.randn(n,1)
    +y = 4+3*x#+np.random.randn(n,1)
     
     X = np.c_[np.ones((n,1)), x]
     XT_X = X.T @ X
    @@ -355,44 +368,32 @@ EigValues, EigVectors = np= np.random.randn(2,1)
     eta = 1.0/np.max(EigValues)
    -Niterations = 100
    +Niterations = 30
     
    -# Note that we request the derivative wrt third argument (theta, 2 here)
    -training_gradient = grad(CostOLS,2)
    +# define the gradient
    +training_gradient = grad(CostOLS)
     
     for iter in range(Niterations):
    -    gradients = (1.0/n)*training_gradient(y, X, theta)
    +    gradients = training_gradient(theta)
         theta -= eta*gradients
    +    print(iter,gradients[0],gradients[1])
     print("theta from own gd")
     print(theta)
     
    -
    -n_epochs = 50
    -M = 5   #size of each minibatch
    -m = int(n/M) #number of minibatches
    -t0, t1 = 5, 50
    -def learning_schedule(t):
    -    return t0/(t+t1)
    -
    -theta = np.random.randn(2,1)
    -
    +# Now improve with momentum gradient descent
     change = 0.0
     delta_momentum = 0.3
    -
    -for epoch in range(n_epochs):
    -    for i in range(m):
    -        random_index = M*np.random.randint(m)
    -        xi = X[random_index:random_index+M]
    -        yi = y[random_index:random_index+M]
    -        gradients = (1.0/M)*training_gradient(yi, xi, theta)
    -        eta = learning_schedule(epoch*m+i)
    -        # calculate update
    -        new_change = eta*gradients+delta_momentum*change
    -        # take a step
    -        theta -= new_change
    -        # save the change
    -        change = new_change
    -print("theta from own sdg with momentum")
    +for iter in range(Niterations):
    +    # calculate gradient
    +    gradients = training_gradient(theta)
    +    # calculate update
    +    new_change = eta*gradients+delta_momentum*change
    +    # take a step
    +    theta -= new_change
    +    # save the change
    +    change = new_change
    +    print(iter,gradients[0],gradients[1])
    +print("theta from own gd wth momentum")
     print(theta)
     
    @@ -435,7 +436,7 @@ delta_momentum = 48
  • 49
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs040.html b/doc/pub/week40/html/._week40-bs040.html index 559edf959..9e3f77813 100644 --- a/doc/pub/week40/html/._week40-bs040.html +++ b/doc/pub/week40/html/._week40-bs040.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,7 +334,8 @@ MathJax.Hub.Config({

     

     

     

    -

    Similar (second order function now) problem but now with AdaGrad

    +

    But noen of these can compete with Newton's method

    +
    @@ -327,54 +343,44 @@ MathJax.Hub.Config({
    -
    # Using Autograd to calculate gradients using AdaGrad and Stochastic Gradient descent
    -# OLS example
    +  
    # Using Newton's method
     from random import random, seed
     import numpy as np
     import autograd.numpy as np
     import matplotlib.pyplot as plt
     from autograd import grad
     
    -# Note change from previous example
    -def CostOLS(y,X,theta):
    -    return np.sum((y-X @ theta)**2)
    +def CostOLS(beta):
    +    return (1.0/n)*np.sum((y-X @ beta)**2)
     
    -n = 1000
    -x = np.random.rand(n,1)
    -y = 2.0+3*x +4*x*x
    +n = 100
    +x = 2*np.random.rand(n,1)
    +y = 4+3*x+np.random.randn(n,1)
     
    -X = np.c_[np.ones((n,1)), x, x*x]
    +X = np.c_[np.ones((n,1)), x]
     XT_X = X.T @ X
    -theta_linreg = np.linalg.pinv(XT_X) @ (X.T @ y)
    +beta_linreg = np.linalg.pinv(XT_X) @ (X.T @ y)
     print("Own inversion")
    -print(theta_linreg)
    +print(beta_linreg)
    +# Hessian matrix
    +H = (2.0/n)* XT_X
    +# Note that here the Hessian does not depend on the parameters beta
    +invH = np.linalg.pinv(H)
    +EigValues, EigVectors = np.linalg.eig(H)
    +print(f"Eigenvalues of Hessian Matrix:{EigValues}")
     
    +beta = np.random.randn(2,1)
    +Niterations = 5
     
    -# Note that we request the derivative wrt third argument (theta, 2 here)
    -training_gradient = grad(CostOLS,2)
    -# Define parameters for Stochastic Gradient Descent
    -n_epochs = 50
    -M = 5   #size of each minibatch
    -m = int(n/M) #number of minibatches
    -# Guess for unknown parameters theta
    -theta = np.random.randn(3,1)
    +# define the gradient
    +training_gradient = grad(CostOLS)
     
    -# Value for learning rate
    -eta = 0.01
    -# Including AdaGrad parameter to avoid possible division by zero
    -delta  = 1e-8
    -for epoch in range(n_epochs):
    -    Giter = 0.0
    -    for i in range(m):
    -        random_index = M*np.random.randint(m)
    -        xi = X[random_index:random_index+M]
    -        yi = y[random_index:random_index+M]
    -        gradients = (1.0/M)*training_gradient(yi, xi, theta)
    -        Giter += gradients*gradients
    -        update = gradients*eta/(delta+np.sqrt(Giter))
    -        theta -= update
    -print("theta from own AdaGrad")
    -print(theta)
    +for iter in range(Niterations):
    +    gradients = training_gradient(beta)
    +    beta -= invH @ gradients
    +    print(iter,gradients[0],gradients[1])
    +print("beta from own Newton code")
    +print(beta)
     
    @@ -390,7 +396,6 @@ delta = 1e-8
    -

    Running this code we note an almost perfect agreement with the results from matrix inversion.

    @@ -417,7 +422,7 @@ delta = 1e-849

  • 50
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs041.html b/doc/pub/week40/html/._week40-bs041.html index 885d97403..c7b952aae 100644 --- a/doc/pub/week40/html/._week40-bs041.html +++ b/doc/pub/week40/html/._week40-bs041.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,7 +334,9 @@ MathJax.Hub.Config({

     

     

     

    -

    RMSprop for adaptive learning rate with Stochastic Gradient Descent

    +

    Including Stochastic Gradient Descent with Autograd

    +

    In this code we include the stochastic gradient descent approach discussed above. Note here that we specify which argument we are taking the derivative with respect to when using autograd.

    +
    @@ -327,7 +344,7 @@ MathJax.Hub.Config({
    -
    # Using Autograd to calculate gradients using RMSprop  and Stochastic Gradient descent
    +  
    # Using Autograd to calculate gradients using SGD
     # OLS example
     from random import random, seed
     import numpy as np
    @@ -339,47 +356,66 @@ MathJax.Hub.Config({
     def CostOLS(y,X,theta):
         return np.sum((y-X @ theta)**2)
     
    -n = 1000
    -x = np.random.rand(n,1)
    -y = 2.0+3*x +4*x*x# +np.random.randn(n,1)
    +n = 100
    +x = 2*np.random.rand(n,1)
    +y = 4+3*x+np.random.randn(n,1)
     
    -X = np.c_[np.ones((n,1)), x, x*x]
    +X = np.c_[np.ones((n,1)), x]
     XT_X = X.T @ X
     theta_linreg = np.linalg.pinv(XT_X) @ (X.T @ y)
     print("Own inversion")
     print(theta_linreg)
    +# Hessian matrix
    +H = (2.0/n)* XT_X
    +EigValues, EigVectors = np.linalg.eig(H)
    +print(f"Eigenvalues of Hessian Matrix:{EigValues}")
     
    +theta = np.random.randn(2,1)
    +eta = 1.0/np.max(EigValues)
    +Niterations = 1000
     
     # Note that we request the derivative wrt third argument (theta, 2 here)
     training_gradient = grad(CostOLS,2)
    -# Define parameters for Stochastic Gradient Descent
    +
    +for iter in range(Niterations):
    +    gradients = (1.0/n)*training_gradient(y, X, theta)
    +    theta -= eta*gradients
    +print("theta from own gd")
    +print(theta)
    +
    +xnew = np.array([[0],[2]])
    +Xnew = np.c_[np.ones((2,1)), xnew]
    +ypredict = Xnew.dot(theta)
    +ypredict2 = Xnew.dot(theta_linreg)
    +
    +plt.plot(xnew, ypredict, "r-")
    +plt.plot(xnew, ypredict2, "b-")
    +plt.plot(x, y ,'ro')
    +plt.axis([0,2.0,0, 15.0])
    +plt.xlabel(r'$x$')
    +plt.ylabel(r'$y$')
    +plt.title(r'Random numbers ')
    +plt.show()
    +
     n_epochs = 50
     M = 5   #size of each minibatch
     m = int(n/M) #number of minibatches
    -# Guess for unknown parameters theta
    -theta = np.random.randn(3,1)
    +t0, t1 = 5, 50
    +def learning_schedule(t):
    +    return t0/(t+t1)
    +
    +theta = np.random.randn(2,1)
     
    -# Value for learning rate
    -eta = 0.01
    -# Value for parameter rho
    -rho = 0.99
    -# Including AdaGrad parameter to avoid possible division by zero
    -delta  = 1e-8
     for epoch in range(n_epochs):
    -    Giter = 0.0
    +# Can you figure out a better way of setting up the contributions to each batch?
         for i in range(m):
             random_index = M*np.random.randint(m)
             xi = X[random_index:random_index+M]
             yi = y[random_index:random_index+M]
             gradients = (1.0/M)*training_gradient(yi, xi, theta)
    -	# Accumulated gradient
    -	# Scaling with rho the new and the previous results
    -        Giter = (rho*Giter+(1-rho)*gradients*gradients)
    -	# Taking the diagonal only and inverting
    -        update = gradients*eta/(delta+np.sqrt(Giter))
    -	# Hadamard product
    -        theta -= update
    -print("theta from own RMSprop")
    +        eta = learning_schedule(epoch*m+i)
    +        theta = theta - eta*gradients
    +print("theta from own sdg")
     print(theta)
     
    @@ -422,7 +458,7 @@ delta = 1e-850
  • 51
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs042.html b/doc/pub/week40/html/._week40-bs042.html index 7db9300b0..0f78f57ba 100644 --- a/doc/pub/week40/html/._week40-bs042.html +++ b/doc/pub/week40/html/._week40-bs042.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,8 +334,7 @@ MathJax.Hub.Config({

     

     

     

    -

    And finally ADAM

    - +

    Same code but now with momentum gradient descent

    @@ -328,7 +342,7 @@ MathJax.Hub.Config({
    -
    # Using Autograd to calculate gradients using RMSprop  and Stochastic Gradient descent
    +  
    # Using Autograd to calculate gradients using SGD
     # OLS example
     from random import random, seed
     import numpy as np
    @@ -340,52 +354,60 @@ MathJax.Hub.Config({
     def CostOLS(y,X,theta):
         return np.sum((y-X @ theta)**2)
     
    -n = 1000
    -x = np.random.rand(n,1)
    -y = 2.0+3*x +4*x*x# +np.random.randn(n,1)
    +n = 100
    +x = 2*np.random.rand(n,1)
    +y = 4+3*x+np.random.randn(n,1)
     
    -X = np.c_[np.ones((n,1)), x, x*x]
    +X = np.c_[np.ones((n,1)), x]
     XT_X = X.T @ X
     theta_linreg = np.linalg.pinv(XT_X) @ (X.T @ y)
     print("Own inversion")
     print(theta_linreg)
    +# Hessian matrix
    +H = (2.0/n)* XT_X
    +EigValues, EigVectors = np.linalg.eig(H)
    +print(f"Eigenvalues of Hessian Matrix:{EigValues}")
     
    +theta = np.random.randn(2,1)
    +eta = 1.0/np.max(EigValues)
    +Niterations = 100
     
     # Note that we request the derivative wrt third argument (theta, 2 here)
     training_gradient = grad(CostOLS,2)
    -# Define parameters for Stochastic Gradient Descent
    +
    +for iter in range(Niterations):
    +    gradients = (1.0/n)*training_gradient(y, X, theta)
    +    theta -= eta*gradients
    +print("theta from own gd")
    +print(theta)
    +
    +
     n_epochs = 50
     M = 5   #size of each minibatch
     m = int(n/M) #number of minibatches
    -# Guess for unknown parameters theta
    -theta = np.random.randn(3,1)
    +t0, t1 = 5, 50
    +def learning_schedule(t):
    +    return t0/(t+t1)
    +
    +theta = np.random.randn(2,1)
    +
    +change = 0.0
    +delta_momentum = 0.3
     
    -# Value for learning rate
    -eta = 0.01
    -# Value for parameters beta1 and beta2, see https://arxiv.org/abs/1412.6980
    -beta1 = 0.9
    -beta2 = 0.999
    -# Including AdaGrad parameter to avoid possible division by zero
    -delta  = 1e-7
    -iter = 0
     for epoch in range(n_epochs):
    -    first_moment = 0.0
    -    second_moment = 0.0
    -    iter += 1
         for i in range(m):
             random_index = M*np.random.randint(m)
             xi = X[random_index:random_index+M]
             yi = y[random_index:random_index+M]
             gradients = (1.0/M)*training_gradient(yi, xi, theta)
    -        # Computing moments first
    -        first_moment = beta1*first_moment + (1-beta1)*gradients
    -        second_moment = beta2*second_moment+(1-beta2)*gradients*gradients
    -        first_term = first_moment/(1.0-beta1**iter)
    -        second_term = second_moment/(1.0-beta2**iter)
    -	# Scaling with rho the new and the previous results
    -        update = eta*first_term/(np.sqrt(second_term)+delta)
    -        theta -= update
    -print("theta from own ADAM")
    +        eta = learning_schedule(epoch*m+i)
    +        # calculate update
    +        new_change = eta*gradients+delta_momentum*change
    +        # take a step
    +        theta -= new_change
    +        # save the change
    +        change = new_change
    +print("theta from own sdg with momentum")
     print(theta)
     
    @@ -428,7 +450,7 @@ delta = 1e-751
  • 52
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs043.html b/doc/pub/week40/html/._week40-bs043.html index 2acb627a0..da9476107 100644 --- a/doc/pub/week40/html/._week40-bs043.html +++ b/doc/pub/week40/html/._week40-bs043.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,8 +334,7 @@ MathJax.Hub.Config({

     

     

     

    -

    And Logistic Regression

    - +

    Similar (second order function now) problem but now with AdaGrad

    @@ -328,39 +342,54 @@ MathJax.Hub.Config({
    -
    import autograd.numpy as np
    +  
    # Using Autograd to calculate gradients using AdaGrad and Stochastic Gradient descent
    +# OLS example
    +from random import random, seed
    +import numpy as np
    +import autograd.numpy as np
    +import matplotlib.pyplot as plt
     from autograd import grad
     
    -def sigmoid(x):
    -    return 0.5 * (np.tanh(x / 2.) + 1)
    +# Note change from previous example
    +def CostOLS(y,X,theta):
    +    return np.sum((y-X @ theta)**2)
     
    -def logistic_predictions(weights, inputs):
    -    # Outputs probability of a label being true according to logistic model.
    -    return sigmoid(np.dot(inputs, weights))
    +n = 1000
    +x = np.random.rand(n,1)
    +y = 2.0+3*x +4*x*x
     
    -def training_loss(weights):
    -    # Training loss is the negative log-likelihood of the training labels.
    -    preds = logistic_predictions(weights, inputs)
    -    label_probabilities = preds * targets + (1 - preds) * (1 - targets)
    -    return -np.sum(np.log(label_probabilities))
    +X = np.c_[np.ones((n,1)), x, x*x]
    +XT_X = X.T @ X
    +theta_linreg = np.linalg.pinv(XT_X) @ (X.T @ y)
    +print("Own inversion")
    +print(theta_linreg)
     
    -# Build a toy dataset.
    -inputs = np.array([[0.52, 1.12,  0.77],
    -                   [0.88, -1.08, 0.15],
    -                   [0.52, 0.06, -1.30],
    -                   [0.74, -2.49, 1.39]])
    -targets = np.array([True, True, False, True])
     
    -# Define a function that returns gradients of training loss using Autograd.
    -training_gradient_fun = grad(training_loss)
    +# Note that we request the derivative wrt third argument (theta, 2 here)
    +training_gradient = grad(CostOLS,2)
    +# Define parameters for Stochastic Gradient Descent
    +n_epochs = 50
    +M = 5   #size of each minibatch
    +m = int(n/M) #number of minibatches
    +# Guess for unknown parameters theta
    +theta = np.random.randn(3,1)
     
    -# Optimize weights using gradient descent.
    -weights = np.array([0.0, 0.0, 0.0])
    -print("Initial loss:", training_loss(weights))
    -for i in range(100):
    -    weights -= training_gradient_fun(weights) * 0.01
    -
    -print("Trained loss:", training_loss(weights))
    +# Value for learning rate
    +eta = 0.01
    +# Including AdaGrad parameter to avoid possible division by zero
    +delta  = 1e-8
    +for epoch in range(n_epochs):
    +    Giter = 0.0
    +    for i in range(m):
    +        random_index = M*np.random.randint(m)
    +        xi = X[random_index:random_index+M]
    +        yi = y[random_index:random_index+M]
    +        gradients = (1.0/M)*training_gradient(yi, xi, theta)
    +        Giter += gradients*gradients
    +        update = gradients*eta/(delta+np.sqrt(Giter))
    +        theta -= update
    +print("theta from own AdaGrad")
    +print(theta)
     
    @@ -376,6 +405,7 @@ weights = np.
    +

    Running this code we note an almost perfect agreement with the results from matrix inversion.

    @@ -402,7 +432,7 @@ weights = np.52

  • 53
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs044.html b/doc/pub/week40/html/._week40-bs044.html index 205a16cbd..192d7a5bf 100644 --- a/doc/pub/week40/html/._week40-bs044.html +++ b/doc/pub/week40/html/._week40-bs044.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,17 +334,7 @@ MathJax.Hub.Config({

     

     

     

    -

    Introducing JAX

    - -

    Presently, instead of using autograd, we recommend using JAX

    - -

    JAX is Autograd and XLA (Accelerated Linear Algebra)), -brought together for high-performance numerical computing and machine learning research. -It provides composable transformations of Python+NumPy programs: differentiate, vectorize, parallelize, Just-In-Time compile to GPU/TPU, and more. -

    - -

    Here's a simple example on how you can use JAX to compute the derivate of the logistic function.

    - +

    RMSprop for adaptive learning rate with Stochastic Gradient Descent

    @@ -337,15 +342,60 @@ It provides composable transformations of Python+NumPy programs: differentiate,
    -
    import jax.numpy as jnp
    -from jax import grad, jit, vmap
    +  
    # Using Autograd to calculate gradients using RMSprop  and Stochastic Gradient descent
    +# OLS example
    +from random import random, seed
    +import numpy as np
    +import autograd.numpy as np
    +import matplotlib.pyplot as plt
    +from autograd import grad
     
    -def sum_logistic(x):
    -  return jnp.sum(1.0 / (1.0 + jnp.exp(-x)))
    +# Note change from previous example
    +def CostOLS(y,X,theta):
    +    return np.sum((y-X @ theta)**2)
     
    -x_small = jnp.arange(3.)
    -derivative_fn = grad(sum_logistic)
    -print(derivative_fn(x_small))
    +n = 1000
    +x = np.random.rand(n,1)
    +y = 2.0+3*x +4*x*x# +np.random.randn(n,1)
    +
    +X = np.c_[np.ones((n,1)), x, x*x]
    +XT_X = X.T @ X
    +theta_linreg = np.linalg.pinv(XT_X) @ (X.T @ y)
    +print("Own inversion")
    +print(theta_linreg)
    +
    +
    +# Note that we request the derivative wrt third argument (theta, 2 here)
    +training_gradient = grad(CostOLS,2)
    +# Define parameters for Stochastic Gradient Descent
    +n_epochs = 50
    +M = 5   #size of each minibatch
    +m = int(n/M) #number of minibatches
    +# Guess for unknown parameters theta
    +theta = np.random.randn(3,1)
    +
    +# Value for learning rate
    +eta = 0.01
    +# Value for parameter rho
    +rho = 0.99
    +# Including AdaGrad parameter to avoid possible division by zero
    +delta  = 1e-8
    +for epoch in range(n_epochs):
    +    Giter = 0.0
    +    for i in range(m):
    +        random_index = M*np.random.randint(m)
    +        xi = X[random_index:random_index+M]
    +        yi = y[random_index:random_index+M]
    +        gradients = (1.0/M)*training_gradient(yi, xi, theta)
    +	# Accumulated gradient
    +	# Scaling with rho the new and the previous results
    +        Giter = (rho*Giter+(1-rho)*gradients*gradients)
    +	# Taking the diagonal only and inverting
    +        update = gradients*eta/(delta+np.sqrt(Giter))
    +	# Hadamard product
    +        theta -= update
    +print("theta from own RMSprop")
    +print(theta)
     
    @@ -387,7 +437,7 @@ derivative_fn = grad(sum_logistic)
  • 53
  • 54
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs045.html b/doc/pub/week40/html/._week40-bs045.html index d00fd1a9d..72316c5e9 100644 --- a/doc/pub/week40/html/._week40-bs045.html +++ b/doc/pub/week40/html/._week40-bs045.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,16 +334,89 @@ MathJax.Hub.Config({

     

     

     

    -

    Introduction to Neural networks

    +

    And finally ADAM

    + + + +
    +
    +
    +
    +
    +
    # Using Autograd to calculate gradients using RMSprop  and Stochastic Gradient descent
    +# OLS example
    +from random import random, seed
    +import numpy as np
    +import autograd.numpy as np
    +import matplotlib.pyplot as plt
    +from autograd import grad
    +
    +# Note change from previous example
    +def CostOLS(y,X,theta):
    +    return np.sum((y-X @ theta)**2)
    +
    +n = 1000
    +x = np.random.rand(n,1)
    +y = 2.0+3*x +4*x*x# +np.random.randn(n,1)
    +
    +X = np.c_[np.ones((n,1)), x, x*x]
    +XT_X = X.T @ X
    +theta_linreg = np.linalg.pinv(XT_X) @ (X.T @ y)
    +print("Own inversion")
    +print(theta_linreg)
    +
    +
    +# Note that we request the derivative wrt third argument (theta, 2 here)
    +training_gradient = grad(CostOLS,2)
    +# Define parameters for Stochastic Gradient Descent
    +n_epochs = 50
    +M = 5   #size of each minibatch
    +m = int(n/M) #number of minibatches
    +# Guess for unknown parameters theta
    +theta = np.random.randn(3,1)
    +
    +# Value for learning rate
    +eta = 0.01
    +# Value for parameters beta1 and beta2, see https://arxiv.org/abs/1412.6980
    +beta1 = 0.9
    +beta2 = 0.999
    +# Including AdaGrad parameter to avoid possible division by zero
    +delta  = 1e-7
    +iter = 0
    +for epoch in range(n_epochs):
    +    first_moment = 0.0
    +    second_moment = 0.0
    +    iter += 1
    +    for i in range(m):
    +        random_index = M*np.random.randint(m)
    +        xi = X[random_index:random_index+M]
    +        yi = y[random_index:random_index+M]
    +        gradients = (1.0/M)*training_gradient(yi, xi, theta)
    +        # Computing moments first
    +        first_moment = beta1*first_moment + (1-beta1)*gradients
    +        second_moment = beta2*second_moment+(1-beta2)*gradients*gradients
    +        first_term = first_moment/(1.0-beta1**iter)
    +        second_term = second_moment/(1.0-beta2**iter)
    +	# Scaling with rho the new and the previous results
    +        update = eta*first_term/(np.sqrt(second_term)+delta)
    +        theta -= update
    +print("theta from own ADAM")
    +print(theta)
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    -

    Artificial neural networks are computational systems that can learn to -perform tasks by considering examples, generally without being -programmed with any task-specific rules. It is supposed to mimic a -biological system, wherein neurons interact by sending signals in the -form of mathematical functions between layers. All layers can contain -an arbitrary number of neurons, and each connection is represented by -a weight variable. -

    @@ -355,7 +443,7 @@ a weight variable.

  • 54
  • 55
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs046.html b/doc/pub/week40/html/._week40-bs046.html index f21db8d22..98dcce6ee 100644 --- a/doc/pub/week40/html/._week40-bs046.html +++ b/doc/pub/week40/html/._week40-bs046.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,64 +334,63 @@ MathJax.Hub.Config({

     

     

     

    -

    Artificial neurons

    +

    And Logistic Regression

    -

    The field of artificial neural networks has a long history of -development, and is closely connected with the advancement of computer -science and computers in general. A model of artificial neurons was -first developed by McCulloch and Pitts in 1943 to study signal -processing in the brain and has later been refined by others. The -general idea is to mimic neural networks in the human brain, which is -composed of billions of neurons that communicate with each other by -sending electrical signals. Each neuron accumulates its incoming -signals, which must exceed an activation threshold to yield an -output. If the threshold is not overcome, the neuron remains inactive, -i.e. has zero output. -

    -

    This behaviour has inspired a simple mathematical model for an artificial neuron.

    + +
    +
    +
    +
    +
    +
    import autograd.numpy as np
    +from autograd import grad
     
    -$$
    -\begin{equation}
    - y = f\left(\sum_{i=1}^n w_ix_i\right) = f(u)
    -\tag{6}
    -\end{equation}
    -$$
    +def sigmoid(x):
    +    return 0.5 * (np.tanh(x / 2.) + 1)
     
    -

    Here, the output \( y \) of the neuron is the value of its activation function, which have as input -a weighted sum of signals \( x_i, \dots ,x_n \) received by \( n \) other neurons. -

    +def logistic_predictions(weights, inputs): + # Outputs probability of a label being true according to logistic model. + return sigmoid(np.dot(inputs, weights)) -

    Conceptually, it is helpful to divide neural networks into four -categories: -

    -
      -
    1. general purpose neural networks for supervised learning,
    2. -
    3. neural networks designed specifically for image processing, the most prominent example of this class being Convolutional Neural Networks (CNNs),
    4. -
    5. neural networks for sequential data such as Recurrent Neural Networks (RNNs), and
    6. -
    7. neural networks for unsupervised learning such as Deep Boltzmann Machines.
    8. -
    -

    In natural science, DNNs and CNNs have already found numerous -applications. In statistical physics, they have been applied to detect -phase transitions in 2D Ising and Potts models, lattice gauge -theories, and different phases of polymers, or solving the -Navier-Stokes equation in weather forecasting. Deep learning has also -found interesting applications in quantum physics. Various quantum -phase transitions can be detected and studied using DNNs and CNNs, -topological phases, and even non-equilibrium many-body -localization. Representing quantum states as DNNs quantum state -tomography are among some of the impressive achievements to reveal the -potential of DNNs to facilitate the study of quantum systems. -

    +def training_loss(weights): + # Training loss is the negative log-likelihood of the training labels. + preds = logistic_predictions(weights, inputs) + label_probabilities = preds * targets + (1 - preds) * (1 - targets) + return -np.sum(np.log(label_probabilities)) -

    In quantum information theory, it has been shown that one can perform -gate decompositions with the help of neural. -

    +# Build a toy dataset. +inputs = np.array([[0.52, 1.12, 0.77], + [0.88, -1.08, 0.15], + [0.52, 0.06, -1.30], + [0.74, -2.49, 1.39]]) +targets = np.array([True, True, False, True]) + +# Define a function that returns gradients of training loss using Autograd. +training_gradient_fun = grad(training_loss) + +# Optimize weights using gradient descent. +weights = np.array([0.0, 0.0, 0.0]) +print("Initial loss:", training_loss(weights)) +for i in range(100): + weights -= training_gradient_fun(weights) * 0.01 + +print("Trained loss:", training_loss(weights)) +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    -

    The applications are not limited to the natural sciences. There is a -plethora of applications in essentially all disciplines, from the -humanities to life science and medicine. -

    @@ -403,7 +417,7 @@ humanities to life science and medicine.

  • 55
  • 56
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs047.html b/doc/pub/week40/html/._week40-bs047.html index 4b2553667..fd45de3d3 100644 --- a/doc/pub/week40/html/._week40-bs047.html +++ b/doc/pub/week40/html/._week40-bs047.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,29 +334,48 @@ MathJax.Hub.Config({

     

     

     

    -

    Neural network types

    +

    Introducing JAX

    -

    An artificial neural network (ANN), is a computational model that -consists of layers of connected neurons, or nodes or units. We will -refer to these interchangeably as units or nodes, and sometimes as -neurons. +

    Presently, instead of using autograd, we recommend using JAX

    + +

    JAX is Autograd and XLA (Accelerated Linear Algebra)), +brought together for high-performance numerical computing and machine learning research. +It provides composable transformations of Python+NumPy programs: differentiate, vectorize, parallelize, Just-In-Time compile to GPU/TPU, and more.

    -

    It is supposed to mimic a biological nervous system by letting each -neuron interact with other neurons by sending signals in the form of -mathematical functions between layers. A wide variety of different -ANNs have been developed, but most of them consist of an input layer, -an output layer and eventual layers in-between, called hidden -layers. All layers can contain an arbitrary number of nodes, and each -connection between two nodes is associated with a weight variable. -

    +

    Here's a simple example on how you can use JAX to compute the derivate of the logistic function.

    + + + +
    +
    +
    +
    +
    +
    import jax.numpy as jnp
    +from jax import grad, jit, vmap
    +
    +def sum_logistic(x):
    +  return jnp.sum(1.0 / (1.0 + jnp.exp(-x)))
    +
    +x_small = jnp.arange(3.)
    +derivative_fn = grad(sum_logistic)
    +print(derivative_fn(x_small))
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    -

    Neural networks (also called neural nets) are neural-inspired -nonlinear models for supervised learning. As we will see, neural nets -can be viewed as natural, more powerful extensions of supervised -learning methods such as linear and logistic regression and soft-max -methods we discussed earlier. -

    @@ -368,7 +402,7 @@ methods we discussed earlier.

  • 56
  • 57
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs048.html b/doc/pub/week40/html/._week40-bs048.html index 6d7b5d165..ed8a3487b 100644 --- a/doc/pub/week40/html/._week40-bs048.html +++ b/doc/pub/week40/html/._week40-bs048.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,19 +334,15 @@ MathJax.Hub.Config({

     

     

     

    -

    Feed-forward neural networks

    +

    Introduction to Neural networks

    -

    The feed-forward neural network (FFNN) was the first and simplest type -of ANNs that were devised. In this network, the information moves in -only one direction: forward through the layers. -

    - -

    Nodes are represented by circles, while the arrows display the -connections between the nodes, including the direction of information -flow. Additionally, each arrow corresponds to a weight variable -(figure to come). We observe that each node in a layer is connected -to all nodes in the subsequent layer, making this a so-called -fully-connected FFNN. +

    Artificial neural networks are computational systems that can learn to +perform tasks by considering examples, generally without being +programmed with any task-specific rules. It is supposed to mimic a +biological system, wherein neurons interact by sending signals in the +form of mathematical functions between layers. All layers can contain +an arbitrary number of neurons, and each connection is represented by +a weight variable.

    @@ -359,7 +370,7 @@ to all nodes in the subsequent layer, making this a so-called

  • 57
  • 58
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs049.html b/doc/pub/week40/html/._week40-bs049.html index 1fec68663..ae3cc7955 100644 --- a/doc/pub/week40/html/._week40-bs049.html +++ b/doc/pub/week40/html/._week40-bs049.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,27 +334,63 @@ MathJax.Hub.Config({

     

     

     

    -

    Convolutional Neural Network

    +

    Artificial neurons

    -

    A different variant of FFNNs are convolutional neural networks -(CNNs), which have a connectivity pattern inspired by the animal -visual cortex. Individual neurons in the visual cortex only respond to -stimuli from small sub-regions of the visual field, called a receptive -field. This makes the neurons well-suited to exploit the strong -spatially local correlation present in natural images. The response of -each neuron can be approximated mathematically as a convolution -operation. (figure to come) +

    The field of artificial neural networks has a long history of +development, and is closely connected with the advancement of computer +science and computers in general. A model of artificial neurons was +first developed by McCulloch and Pitts in 1943 to study signal +processing in the brain and has later been refined by others. The +general idea is to mimic neural networks in the human brain, which is +composed of billions of neurons that communicate with each other by +sending electrical signals. Each neuron accumulates its incoming +signals, which must exceed an activation threshold to yield an +output. If the threshold is not overcome, the neuron remains inactive, +i.e. has zero output.

    -

    Convolutional neural networks emulate the behaviour of neurons in the -visual cortex by enforcing a local connectivity pattern between -nodes of adjacent layers: Each node in a convolutional layer is -connected only to a subset of the nodes in the previous layer, in -contrast to the fully-connected FFNN. Often, CNNs consist of several -convolutional layers that learn local features of the input, with a -fully-connected layer at the end, which gathers all the local data and -produces the outputs. They have wide applications in image and video -recognition. +

    This behaviour has inspired a simple mathematical model for an artificial neuron.

    + +$$ +\begin{equation} + y = f\left(\sum_{i=1}^n w_ix_i\right) = f(u) +\tag{6} +\end{equation} +$$ + +

    Here, the output \( y \) of the neuron is the value of its activation function, which have as input +a weighted sum of signals \( x_i, \dots ,x_n \) received by \( n \) other neurons. +

    + +

    Conceptually, it is helpful to divide neural networks into four +categories: +

    +
      +
    1. general purpose neural networks for supervised learning,
    2. +
    3. neural networks designed specifically for image processing, the most prominent example of this class being Convolutional Neural Networks (CNNs),
    4. +
    5. neural networks for sequential data such as Recurrent Neural Networks (RNNs), and
    6. +
    7. neural networks for unsupervised learning such as Deep Boltzmann Machines.
    8. +
    +

    In natural science, DNNs and CNNs have already found numerous +applications. In statistical physics, they have been applied to detect +phase transitions in 2D Ising and Potts models, lattice gauge +theories, and different phases of polymers, or solving the +Navier-Stokes equation in weather forecasting. Deep learning has also +found interesting applications in quantum physics. Various quantum +phase transitions can be detected and studied using DNNs and CNNs, +topological phases, and even non-equilibrium many-body +localization. Representing quantum states as DNNs quantum state +tomography are among some of the impressive achievements to reveal the +potential of DNNs to facilitate the study of quantum systems. +

    + +

    In quantum information theory, it has been shown that one can perform +gate decompositions with the help of neural. +

    + +

    The applications are not limited to the natural sciences. There is a +plethora of applications in essentially all disciplines, from the +humanities to life science and medicine.

    @@ -367,7 +418,7 @@ recognition.

  • 58
  • 59
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs050.html b/doc/pub/week40/html/._week40-bs050.html index a6fd6363c..d4a172f99 100644 --- a/doc/pub/week40/html/._week40-bs050.html +++ b/doc/pub/week40/html/._week40-bs050.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,18 +334,28 @@ MathJax.Hub.Config({

     

     

     

    -

    Recurrent neural networks

    +

    Neural network types

    -

    So far we have only mentioned ANNs where information flows in one -direction: forward. Recurrent neural networks on the other hand, -have connections between nodes that form directed cycles. This -creates a form of internal memory which are able to capture -information on what has been calculated before; the output is -dependent on the previous computations. Recurrent NNs make use of -sequential information by performing the same task for every element -in a sequence, where each element depends on previous elements. An -example of such information is sentences, making recurrent NNs -especially well-suited for handwriting and speech recognition. +

    An artificial neural network (ANN), is a computational model that +consists of layers of connected neurons, or nodes or units. We will +refer to these interchangeably as units or nodes, and sometimes as +neurons. +

    + +

    It is supposed to mimic a biological nervous system by letting each +neuron interact with other neurons by sending signals in the form of +mathematical functions between layers. A wide variety of different +ANNs have been developed, but most of them consist of an input layer, +an output layer and eventual layers in-between, called hidden +layers. All layers can contain an arbitrary number of nodes, and each +connection between two nodes is associated with a weight variable. +

    + +

    Neural networks (also called neural nets) are neural-inspired +nonlinear models for supervised learning. As we will see, neural nets +can be viewed as natural, more powerful extensions of supervised +learning methods such as linear and logistic regression and soft-max +methods we discussed earlier.

    @@ -358,7 +383,7 @@ especially well-suited for handwriting and speech recognition.

  • 59
  • 60
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs051.html b/doc/pub/week40/html/._week40-bs051.html index 8d5d0b948..f44d88736 100644 --- a/doc/pub/week40/html/._week40-bs051.html +++ b/doc/pub/week40/html/._week40-bs051.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,18 +334,19 @@ MathJax.Hub.Config({

     

     

     

    -

    Other types of networks

    +

    Feed-forward neural networks

    -

    There are many other kinds of ANNs that have been developed. One type -that is specifically designed for interpolation in multidimensional -space is the radial basis function (RBF) network. RBFs are typically -made up of three layers: an input layer, a hidden layer with -non-linear radial symmetric activation functions and a linear output -layer (''linear'' here means that each node in the output layer has a -linear activation function). The layers are normally fully-connected -and there are no cycles, thus RBFs can be viewed as a type of -fully-connected FFNN. They are however usually treated as a separate -type of NN due the unusual activation functions. +

    The feed-forward neural network (FFNN) was the first and simplest type +of ANNs that were devised. In this network, the information moves in +only one direction: forward through the layers. +

    + +

    Nodes are represented by circles, while the arrows display the +connections between the nodes, including the direction of information +flow. Additionally, each arrow corresponds to a weight variable +(figure to come). We observe that each node in a layer is connected +to all nodes in the subsequent layer, making this a so-called +fully-connected FFNN.

    @@ -358,7 +374,7 @@ type of NN due the unusual activation functions.

  • 60
  • 61
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs052.html b/doc/pub/week40/html/._week40-bs052.html index 718546bcc..255945f6e 100644 --- a/doc/pub/week40/html/._week40-bs052.html +++ b/doc/pub/week40/html/._week40-bs052.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,15 +334,28 @@ MathJax.Hub.Config({

     

     

     

    -

    Multilayer perceptrons

    +

    Convolutional Neural Network

    -

    One uses often so-called fully-connected feed-forward neural networks -with three or more layers (an input layer, one or more hidden layers -and an output layer) consisting of neurons that have non-linear -activation functions. +

    A different variant of FFNNs are convolutional neural networks +(CNNs), which have a connectivity pattern inspired by the animal +visual cortex. Individual neurons in the visual cortex only respond to +stimuli from small sub-regions of the visual field, called a receptive +field. This makes the neurons well-suited to exploit the strong +spatially local correlation present in natural images. The response of +each neuron can be approximated mathematically as a convolution +operation. (figure to come)

    -

    Such networks are often called multilayer perceptrons (MLPs).

    +

    Convolutional neural networks emulate the behaviour of neurons in the +visual cortex by enforcing a local connectivity pattern between +nodes of adjacent layers: Each node in a convolutional layer is +connected only to a subset of the nodes in the previous layer, in +contrast to the fully-connected FFNN. Often, CNNs consist of several +convolutional layers that learn local features of the input, with a +fully-connected layer at the end, which gathers all the local data and +produces the outputs. They have wide applications in image and video +recognition. +

    @@ -354,7 +382,7 @@ activation functions.

  • 61
  • 62
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs053.html b/doc/pub/week40/html/._week40-bs053.html index a562339a2..7d61db15b 100644 --- a/doc/pub/week40/html/._week40-bs053.html +++ b/doc/pub/week40/html/._week40-bs053.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,19 +334,18 @@ MathJax.Hub.Config({

     

     

     

    -

    Why multilayer perceptrons?

    +

    Recurrent neural networks

    -

    According to the Universal approximation theorem, a feed-forward -neural network with just a single hidden layer containing a finite -number of neurons can approximate a continuous multidimensional -function to arbitrary accuracy, assuming the activation function for -the hidden layer is a non-constant, bounded and -monotonically-increasing continuous function. -

    - -

    Note that the requirements on the activation function only applies to -the hidden layer, the output nodes are always assumed to be linear, so -as to not restrict the range of output values. +

    So far we have only mentioned ANNs where information flows in one +direction: forward. Recurrent neural networks on the other hand, +have connections between nodes that form directed cycles. This +creates a form of internal memory which are able to capture +information on what has been calculated before; the output is +dependent on the previous computations. Recurrent NNs make use of +sequential information by performing the same task for every element +in a sequence, where each element depends on previous elements. An +example of such information is sentences, making recurrent NNs +especially well-suited for handwriting and speech recognition.

    @@ -359,7 +373,7 @@ as to not restrict the range of output values.

  • 62
  • 63
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs054.html b/doc/pub/week40/html/._week40-bs054.html index 1a944b2f6..fb09f492f 100644 --- a/doc/pub/week40/html/._week40-bs054.html +++ b/doc/pub/week40/html/._week40-bs054.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,15 +334,19 @@ MathJax.Hub.Config({

     

     

     

    -

    Illustration of a single perceptron model and a multi-perceptron model

    +

    Other types of networks

    -
    -
    -
    -

    Figure 1: In a) we show a single perceptron model while in b) we dispay a network with two hidden layers, an input layer and an output layer.

    -
    -

    -
    +

    There are many other kinds of ANNs that have been developed. One type +that is specifically designed for interpolation in multidimensional +space is the radial basis function (RBF) network. RBFs are typically +made up of three layers: an input layer, a hidden layer with +non-linear radial symmetric activation functions and a linear output +layer (''linear'' here means that each node in the output layer has a +linear activation function). The layers are normally fully-connected +and there are no cycles, thus RBFs can be viewed as a type of +fully-connected FFNN. They are however usually treated as a separate +type of NN due the unusual activation functions. +

    @@ -354,7 +373,7 @@ MathJax.Hub.Config({

  • 63
  • 64
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs055.html b/doc/pub/week40/html/._week40-bs055.html index e520ff00a..fb92505d1 100644 --- a/doc/pub/week40/html/._week40-bs055.html +++ b/doc/pub/week40/html/._week40-bs055.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,68 +334,15 @@ MathJax.Hub.Config({

     

     

     

    -

    Examples of XOR, OR and AND gates

    +

    Multilayer perceptrons

    -

    Let us first try to fit various gates using standard linear -regression. The gates we are thinking of are the classical XOR, OR and -AND gates, well-known elements in computer science. The tables here -show how we can set up the inputs \( x_1 \) and \( x_2 \) in order to yield a -specific target \( y_i \). +

    One uses often so-called fully-connected feed-forward neural networks +with three or more layers (an input layer, one or more hidden layers +and an output layer) consisting of neurons that have non-linear +activation functions.

    - - -
    -
    -
    -
    -
    -
    """
    -Simple code that tests XOR, OR and AND gates with linear regression
    -"""
    -
    -import numpy as np
    -# Design matrix
    -X = np.array([ [1, 0, 0], [1, 0, 1], [1, 1, 0],[1, 1, 1]],dtype=np.float64)
    -print(f"The X.TX  matrix:{X.T @ X}")
    -Xinv = np.linalg.pinv(X.T @ X)
    -print(f"The invers of X.TX  matrix:{Xinv}")
    -
    -# The XOR gate 
    -yXOR = np.array( [ 0, 1 ,1, 0])
    -ThetaXOR  = Xinv @ X.T @ yXOR
    -print(f"The values of theta for the XOR gate:{ThetaXOR}")
    -print(f"The linear regression prediction  for the XOR gate:{X @ ThetaXOR}")
    -
    -
    -# The OR gate 
    -yOR = np.array( [ 0, 1 ,1, 1])
    -ThetaOR  = Xinv @ X.T @ yOR
    -print(f"The values of theta for the OR gate:{ThetaOR}")
    -print(f"The linear regression prediction  for the OR gate:{X @ ThetaOR}")
    -
    -
    -# The OR gate 
    -yAND = np.array( [ 0, 0 ,0, 1])
    -ThetaAND  = Xinv @ X.T @ yAND
    -print(f"The values of theta for the AND gate:{ThetaAND}")
    -print(f"The linear regression prediction  for the AND gate:{X @ ThetaAND}")
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - -

    What is happening here?

    +

    Such networks are often called multilayer perceptrons (MLPs).

    @@ -407,7 +369,7 @@ ThetaAND = Xinv 64

  • 65
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs056.html b/doc/pub/week40/html/._week40-bs056.html index 0cfa3e896..ab32bd66e 100644 --- a/doc/pub/week40/html/._week40-bs056.html +++ b/doc/pub/week40/html/._week40-bs056.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,80 +334,20 @@ MathJax.Hub.Config({

     

     

     

    -

    Does Logistic Regression do a better Job?

    +

    Why multilayer perceptrons?

    +

    According to the Universal approximation theorem, a feed-forward +neural network with just a single hidden layer containing a finite +number of neurons can approximate a continuous multidimensional +function to arbitrary accuracy, assuming the activation function for +the hidden layer is a non-constant, bounded and +monotonically-increasing continuous function. +

    - -
    -
    -
    -
    -
    -
    """
    -Simple code that tests XOR and OR gates with linear regression
    -and logistic regression
    -"""
    -
    -import matplotlib.pyplot as plt
    -from sklearn.linear_model import LogisticRegression
    -import numpy as np
    -
    -# Design matrix
    -X = np.array([ [1, 0, 0], [1, 0, 1], [1, 1, 0],[1, 1, 1]],dtype=np.float64)
    -print(f"The X.TX  matrix:{X.T @ X}")
    -Xinv = np.linalg.pinv(X.T @ X)
    -print(f"The invers of X.TX  matrix:{Xinv}")
    -
    -# The XOR gate 
    -yXOR = np.array( [ 0, 1 ,1, 0])
    -ThetaXOR  = Xinv @ X.T @ yXOR
    -print(f"The values of theta for the XOR gate:{ThetaXOR}")
    -print(f"The linear regression prediction  for the XOR gate:{X @ ThetaXOR}")
    -
    -
    -# The OR gate 
    -yOR = np.array( [ 0, 1 ,1, 1])
    -ThetaOR  = Xinv @ X.T @ yOR
    -print(f"The values of theta for the OR gate:{ThetaOR}")
    -print(f"The linear regression prediction  for the OR gate:{X @ ThetaOR}")
    -
    -
    -# The OR gate 
    -yAND = np.array( [ 0, 0 ,0, 1])
    -ThetaAND  = Xinv @ X.T @ yAND
    -print(f"The values of theta for the AND gate:{ThetaAND}")
    -print(f"The linear regression prediction  for the AND gate:{X @ ThetaAND}")
    -
    -# Now we change to logistic regression
    -
    -
    -# Logistic Regression
    -logreg = LogisticRegression()
    -logreg.fit(X, yOR)
    -print("Test set accuracy with Logistic Regression for OR gate: {:.2f}".format(logreg.score(X,yOR)))
    -
    -logreg.fit(X, yXOR)
    -print("Test set accuracy with Logistic Regression for XOR gate: {:.2f}".format(logreg.score(X,yXOR)))
    -
    -
    -logreg.fit(X, yAND)
    -print("Test set accuracy with Logistic Regression for AND gate: {:.2f}".format(logreg.score(X,yAND)))
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    - -

    Not exactly impressive, but somewhat better.

    +

    Note that the requirements on the activation function only applies to +the hidden layer, the output nodes are always assumed to be linear, so +as to not restrict the range of output values. +

    @@ -419,7 +374,7 @@ logreg.fit(X, yAND)

  • 65
  • 66
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs057.html b/doc/pub/week40/html/._week40-bs057.html index 5eacabe04..1bda10a4f 100644 --- a/doc/pub/week40/html/._week40-bs057.html +++ b/doc/pub/week40/html/._week40-bs057.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,38 +334,15 @@ MathJax.Hub.Config({

     

     

     

    -

    Adding Neural Networks

    - - - -
    -
    -
    -
    -
    -
    # and now neural networks with Scikit-Learn and the XOR
    -
    -from sklearn.neural_network import MLPClassifier
    -from sklearn.datasets import make_classification
    -X, yXOR = make_classification(n_samples=100, random_state=1)
    -FFNN = MLPClassifier(random_state=1, max_iter=300).fit(X, yXOR)
    -FFNN.predict_proba(X)
    -print(f"Test set accuracy with Feed Forward Neural Network  for XOR gate:{FFNN.score(X, yXOR)}")
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    -
    +

    Illustration of a single perceptron model and a multi-perceptron model

    +
    +
    +
    +

    Figure 1: In a) we show a single perceptron model while in b) we dispay a network with two hidden layers, an input layer and an output layer.

    +
    +

    +

    @@ -377,7 +369,7 @@ FFNN.predict_proba(X)

  • 66
  • 67
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs058.html b/doc/pub/week40/html/._week40-bs058.html index 3c181b868..d369936d0 100644 --- a/doc/pub/week40/html/._week40-bs058.html +++ b/doc/pub/week40/html/._week40-bs058.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,21 +334,69 @@ MathJax.Hub.Config({

     

     

     

    -

    Mathematical model

    +

    Examples of XOR, OR and AND gates

    -

    The output \( y \) is produced via the activation function \( f \)

    -$$ - y = f\left(\sum_{i=1}^n w_ix_i + b_i\right) = f(z), -$$ - -

    This function receives \( x_i \) as inputs. -Here the activation \( z=(\sum_{i=1}^n w_ix_i+b_i) \). -In an FFNN of such neurons, the inputs \( x_i \) are the outputs of -the neurons in the preceding layer. Furthermore, an MLP is -fully-connected, which means that each neuron receives a weighted sum -of the outputs of all neurons in the previous layer. +

    Let us first try to fit various gates using standard linear +regression. The gates we are thinking of are the classical XOR, OR and +AND gates, well-known elements in computer science. The tables here +show how we can set up the inputs \( x_1 \) and \( x_2 \) in order to yield a +specific target \( y_i \).

    + + +
    +
    +
    +
    +
    +
    """
    +Simple code that tests XOR, OR and AND gates with linear regression
    +"""
    +
    +import numpy as np
    +# Design matrix
    +X = np.array([ [1, 0, 0], [1, 0, 1], [1, 1, 0],[1, 1, 1]],dtype=np.float64)
    +print(f"The X.TX  matrix:{X.T @ X}")
    +Xinv = np.linalg.pinv(X.T @ X)
    +print(f"The invers of X.TX  matrix:{Xinv}")
    +
    +# The XOR gate 
    +yXOR = np.array( [ 0, 1 ,1, 0])
    +ThetaXOR  = Xinv @ X.T @ yXOR
    +print(f"The values of theta for the XOR gate:{ThetaXOR}")
    +print(f"The linear regression prediction  for the XOR gate:{X @ ThetaXOR}")
    +
    +
    +# The OR gate 
    +yOR = np.array( [ 0, 1 ,1, 1])
    +ThetaOR  = Xinv @ X.T @ yOR
    +print(f"The values of theta for the OR gate:{ThetaOR}")
    +print(f"The linear regression prediction  for the OR gate:{X @ ThetaOR}")
    +
    +
    +# The OR gate 
    +yAND = np.array( [ 0, 0 ,0, 1])
    +ThetaAND  = Xinv @ X.T @ yAND
    +print(f"The values of theta for the AND gate:{ThetaAND}")
    +print(f"The linear regression prediction  for the AND gate:{X @ ThetaAND}")
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    + +

    What is happening here?

    +

      @@ -358,6 +421,8 @@ of the outputs of all neurons in the previous layer.
    • 66
    • 67
    • 68
    • +
    • ...
    • +
    • 71
    • »
    diff --git a/doc/pub/week40/html/._week40-bs059.html b/doc/pub/week40/html/._week40-bs059.html index d1b8d3ab2..3e0f4f2fe 100644 --- a/doc/pub/week40/html/._week40-bs059.html +++ b/doc/pub/week40/html/._week40-bs059.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,48 +334,80 @@ MathJax.Hub.Config({

     

     

     

    -

    Mathematical model

    +

    Does Logistic Regression do a better Job?

    -

    First, for each node \( i \) in the first hidden layer, we calculate a weighted sum \( z_i^1 \) of the input coordinates \( x_j \),

    -$$ -\begin{equation} z_i^1 = \sum_{j=1}^{M} w_{ij}^1 x_j + b_i^1 -\tag{7} -\end{equation} -$$ + +
    +
    +
    +
    +
    +
    """
    +Simple code that tests XOR and OR gates with linear regression
    +and logistic regression
    +"""
     
    -

    Here \( b_i \) is the so-called bias which is normally needed in -case of zero activation weights or inputs. How to fix the biases and -the weights will be discussed below. The value of \( z_i^1 \) is the -argument to the activation function \( f_i \) of each node \( i \), The -variable \( M \) stands for all possible inputs to a given node \( i \) in the -first layer. We define the output \( y_i^1 \) of all neurons in layer 1 as -

    +import matplotlib.pyplot as plt +from sklearn.linear_model import LogisticRegression +import numpy as np -$$ -\begin{equation} - y_i^1 = f(z_i^1) = f\left(\sum_{j=1}^M w_{ij}^1 x_j + b_i^1\right) -\tag{8} -\end{equation} -$$ +# Design matrix +X = np.array([ [1, 0, 0], [1, 0, 1], [1, 1, 0],[1, 1, 1]],dtype=np.float64) +print(f"The X.TX matrix:{X.T @ X}") +Xinv = np.linalg.pinv(X.T @ X) +print(f"The invers of X.TX matrix:{Xinv}") -

    where we assume that all nodes in the same layer have identical -activation functions, hence the notation \( f \). In general, we could assume in the more general case that different layers have different activation functions. -In this case we would identify these functions with a superscript \( l \) for the \( l \)-th layer, -

    +# The XOR gate +yXOR = np.array( [ 0, 1 ,1, 0]) +ThetaXOR = Xinv @ X.T @ yXOR +print(f"The values of theta for the XOR gate:{ThetaXOR}") +print(f"The linear regression prediction for the XOR gate:{X @ ThetaXOR}") -$$ -\begin{equation} - y_i^l = f^l(u_i^l) = f^l\left(\sum_{j=1}^{N_{l-1}} w_{ij}^l y_j^{l-1} + b_i^l\right) -\tag{9} -\end{equation} -$$ -

    where \( N_l \) is the number of nodes in layer \( l \). When the output of -all the nodes in the first hidden layer are computed, the values of -the subsequent layer can be calculated and so forth until the output -is obtained. -

    +# The OR gate +yOR = np.array( [ 0, 1 ,1, 1]) +ThetaOR = Xinv @ X.T @ yOR +print(f"The values of theta for the OR gate:{ThetaOR}") +print(f"The linear regression prediction for the OR gate:{X @ ThetaOR}") + + +# The OR gate +yAND = np.array( [ 0, 0 ,0, 1]) +ThetaAND = Xinv @ X.T @ yAND +print(f"The values of theta for the AND gate:{ThetaAND}") +print(f"The linear regression prediction for the AND gate:{X @ ThetaAND}") + +# Now we change to logistic regression + + +# Logistic Regression +logreg = LogisticRegression() +logreg.fit(X, yOR) +print("Test set accuracy with Logistic Regression for OR gate: {:.2f}".format(logreg.score(X,yOR))) + +logreg.fit(X, yXOR) +print("Test set accuracy with Logistic Regression for XOR gate: {:.2f}".format(logreg.score(X,yXOR))) + + +logreg.fit(X, yAND) +print("Test set accuracy with Logistic Regression for AND gate: {:.2f}".format(logreg.score(X,yAND))) +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    + +

    Not exactly impressive, but somewhat better.

    @@ -385,6 +432,9 @@ is obtained.

  • 66
  • 67
  • 68
  • +
  • 69
  • +
  • ...
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs060.html b/doc/pub/week40/html/._week40-bs060.html index e2a7b7560..6380c144e 100644 --- a/doc/pub/week40/html/._week40-bs060.html +++ b/doc/pub/week40/html/._week40-bs060.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,30 +334,37 @@ MathJax.Hub.Config({

     

     

     

    -

    Mathematical model

    +

    Adding Neural Networks

    -

    The output of neuron \( i \) in layer 2 is thus,

    -$$ -\begin{align} - y_i^2 &= f^2\left(\sum_{j=1}^N w_{ij}^2 y_j^1 + b_i^2\right) -\tag{10}\\ - &= f^2\left[\sum_{j=1}^N w_{ij}^2f^1\left(\sum_{k=1}^M w_{jk}^1 x_k + b_j^1\right) + b_i^2\right] -\tag{11} -\end{align} -$$ + +
    +
    +
    +
    +
    +
    # and now neural networks with Scikit-Learn and the XOR
     
    -

    where we have substituted \( y_k^1 \) with the inputs \( x_k \). Finally, the ANN output reads

    - -$$ -\begin{align} - y_i^3 &= f^3\left(\sum_{j=1}^N w_{ij}^3 y_j^2 + b_i^3\right) -\tag{12}\\ - &= f_3\left[\sum_{j} w_{ij}^3 f^2\left(\sum_{k} w_{jk}^2 f^1\left(\sum_{m} w_{km}^1 x_m + b_k^1\right) + b_j^2\right) - + b_1^3\right] -\tag{13} -\end{align} -$$ +from sklearn.neural_network import MLPClassifier +from sklearn.datasets import make_classification +X, yXOR = make_classification(n_samples=100, random_state=1) +FFNN = MLPClassifier(random_state=1, max_iter=300).fit(X, yXOR) +FFNN.predict_proba(X) +print(f"Test set accuracy with Feed Forward Neural Network for XOR gate:{FFNN.score(X, yXOR)}") +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +
    +

    @@ -367,6 +389,10 @@ $$

  • 66
  • 67
  • 68
  • +
  • 69
  • +
  • 70
  • +
  • ...
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs061.html b/doc/pub/week40/html/._week40-bs061.html index b8d53fc97..faca65769 100644 --- a/doc/pub/week40/html/._week40-bs061.html +++ b/doc/pub/week40/html/._week40-bs061.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -321,19 +336,17 @@ MathJax.Hub.Config({

    Mathematical model

    -

    We can generalize this expression to an MLP with \( l \) hidden -layers. The complete functional form is, -

    - +

    The output \( y \) is produced via the activation function \( f \)

    $$ -\begin{align} -&y^{l+1}_i = f^{l+1}\left[\!\sum_{j=1}^{N_l} w_{ij}^3 f^l\left(\sum_{k=1}^{N_{l-1}}w_{jk}^{l-1}\left(\dots f^1\left(\sum_{n=1}^{N_0} w_{mn}^1 x_n+ b_m^1\right)\dots\right)+b_k^2\right)+b_1^3\right] && -\tag{14} -\end{align} + y = f\left(\sum_{i=1}^n w_ix_i + b_i\right) = f(z), $$ -

    which illustrates a basic property of MLPs: The only independent -variables are the input values \( x_n \). +

    This function receives \( x_i \) as inputs. +Here the activation \( z=(\sum_{i=1}^n w_ix_i+b_i) \). +In an FFNN of such neurons, the inputs \( x_i \) are the outputs of +the neurons in the preceding layer. Furthermore, an MLP is +fully-connected, which means that each neuron receives a weighted sum +of the outputs of all neurons in the previous layer.

    @@ -357,6 +370,9 @@ variables are the input values \( x_n \).

  • 66
  • 67
  • 68
  • +
  • 69
  • +
  • 70
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs062.html b/doc/pub/week40/html/._week40-bs062.html index e5089c37b..5a371cf39 100644 --- a/doc/pub/week40/html/._week40-bs062.html +++ b/doc/pub/week40/html/._week40-bs062.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -321,28 +336,45 @@ MathJax.Hub.Config({

    Mathematical model

    -

    This confirms that an MLP, despite its quite convoluted mathematical -form, is nothing more than an analytic function, specifically a -mapping of real-valued vectors \( \hat{x} \in \mathbb{R}^n \rightarrow -\hat{y} \in \mathbb{R}^m \). -

    +

    First, for each node \( i \) in the first hidden layer, we calculate a weighted sum \( z_i^1 \) of the input coordinates \( x_j \),

    -

    Furthermore, the flexibility and universality of an MLP can be -illustrated by realizing that the expression is essentially a nested -sum of scaled activation functions of the form +$$ +\begin{equation} z_i^1 = \sum_{j=1}^{M} w_{ij}^1 x_j + b_i^1 +\tag{7} +\end{equation} +$$ + +

    Here \( b_i \) is the so-called bias which is normally needed in +case of zero activation weights or inputs. How to fix the biases and +the weights will be discussed below. The value of \( z_i^1 \) is the +argument to the activation function \( f_i \) of each node \( i \), The +variable \( M \) stands for all possible inputs to a given node \( i \) in the +first layer. We define the output \( y_i^1 \) of all neurons in layer 1 as

    $$ \begin{equation} - f(x) = c_1 f(c_2 x + c_3) + c_4 -\tag{15} + y_i^1 = f(z_i^1) = f\left(\sum_{j=1}^M w_{ij}^1 x_j + b_i^1\right) +\tag{8} \end{equation} $$ -

    where the parameters \( c_i \) are weights and biases. By adjusting these -parameters, the activation functions can be shifted up and down or -left and right, change slope or be rescaled which is the key to the -flexibility of a neural network. +

    where we assume that all nodes in the same layer have identical +activation functions, hence the notation \( f \). In general, we could assume in the more general case that different layers have different activation functions. +In this case we would identify these functions with a superscript \( l \) for the \( l \)-th layer, +

    + +$$ +\begin{equation} + y_i^l = f^l(u_i^l) = f^l\left(\sum_{j=1}^{N_{l-1}} w_{ij}^l y_j^{l-1} + b_i^l\right) +\tag{9} +\end{equation} +$$ + +

    where \( N_l \) is the number of nodes in layer \( l \). When the output of +all the nodes in the first hidden layer are computed, the values of +the subsequent layer can be calculated and so forth until the output +is obtained.

    @@ -365,6 +397,9 @@ flexibility of a neural network.

  • 66
  • 67
  • 68
  • +
  • 69
  • +
  • 70
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs063.html b/doc/pub/week40/html/._week40-bs063.html index 160f971d8..7f48aecca 100644 --- a/doc/pub/week40/html/._week40-bs063.html +++ b/doc/pub/week40/html/._week40-bs063.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,39 +334,29 @@ MathJax.Hub.Config({

     

     

     

    -

    Matrix-vector notation

    +

    Mathematical model

    -

    We can introduce a more convenient notation for the activations in an A NN.

    +

    The output of neuron \( i \) in layer 2 is thus,

    -

    Additionally, we can represent the biases and activations -as layer-wise column vectors \( \hat{b}_l \) and \( \hat{y}_l \), so that the \( i \)-th element of each vector -is the bias \( b_i^l \) and activation \( y_i^l \) of node \( i \) in layer \( l \) respectively. -

    - -

    We have that \( \mathrm{W}_l \) is an \( N_{l-1} \times N_l \) matrix, while \( \hat{b}_l \) and \( \hat{y}_l \) are \( N_l \times 1 \) column vectors. -With this notation, the sum becomes a matrix-vector multiplication, and we can write -the equation for the activations of hidden layer 2 (assuming three nodes for simplicity) as -

    $$ -\begin{equation} - \hat{y}_2 = f_2(\mathrm{W}_2 \hat{y}_{1} + \hat{b}_{2}) = - f_2\left(\left[\begin{array}{ccc} - w^2_{11} &w^2_{12} &w^2_{13} \\ - w^2_{21} &w^2_{22} &w^2_{23} \\ - w^2_{31} &w^2_{32} &w^2_{33} \\ - \end{array} \right] \cdot - \left[\begin{array}{c} - y^1_1 \\ - y^1_2 \\ - y^1_3 \\ - \end{array}\right] + - \left[\begin{array}{c} - b^2_1 \\ - b^2_2 \\ - b^2_3 \\ - \end{array}\right]\right). -\tag{16} -\end{equation} +\begin{align} + y_i^2 &= f^2\left(\sum_{j=1}^N w_{ij}^2 y_j^1 + b_i^2\right) +\tag{10}\\ + &= f^2\left[\sum_{j=1}^N w_{ij}^2f^1\left(\sum_{k=1}^M w_{jk}^1 x_k + b_j^1\right) + b_i^2\right] +\tag{11} +\end{align} +$$ + +

    where we have substituted \( y_k^1 \) with the inputs \( x_k \). Finally, the ANN output reads

    + +$$ +\begin{align} + y_i^3 &= f^3\left(\sum_{j=1}^N w_{ij}^3 y_j^2 + b_i^3\right) +\tag{12}\\ + &= f_3\left[\sum_{j} w_{ij}^3 f^2\left(\sum_{k} w_{jk}^2 f^1\left(\sum_{m} w_{km}^1 x_m + b_k^1\right) + b_j^2\right) + + b_1^3\right] +\tag{13} +\end{align} $$ @@ -374,6 +379,9 @@ $$
  • 66
  • 67
  • 68
  • +
  • 69
  • +
  • 70
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs064.html b/doc/pub/week40/html/._week40-bs064.html index 725401530..dfdcc07a6 100644 --- a/doc/pub/week40/html/._week40-bs064.html +++ b/doc/pub/week40/html/._week40-bs064.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,23 +334,21 @@ MathJax.Hub.Config({

     

     

     

    -

    Matrix-vector notation and activation

    +

    Mathematical model

    -

    The activation of node \( i \) in layer 2 is

    +

    We can generalize this expression to an MLP with \( l \) hidden +layers. The complete functional form is, +

    $$ -\begin{equation} - y^2_i = f_2\Bigr(w^2_{i1}y^1_1 + w^2_{i2}y^1_2 + w^2_{i3}y^1_3 + b^2_i\Bigr) = - f_2\left(\sum_{j=1}^3 w^2_{ij} y_j^1 + b^2_i\right). -\tag{17} -\end{equation} +\begin{align} +&y^{l+1}_i = f^{l+1}\left[\!\sum_{j=1}^{N_l} w_{ij}^3 f^l\left(\sum_{k=1}^{N_{l-1}}w_{jk}^{l-1}\left(\dots f^1\left(\sum_{n=1}^{N_0} w_{mn}^1 x_n+ b_m^1\right)\dots\right)+b_k^2\right)+b_1^3\right] && +\tag{14} +\end{align} $$ -

    This is not just a convenient and compact notation, but also a useful -and intuitive way to think about MLPs: The output is calculated by a -series of matrix-vector multiplications and vector additions that are -used as input to the activation functions. For each operation -\( \mathrm{W}_l \hat{y}_{l-1} \) we move forward one layer. +

    which illustrates a basic property of MLPs: The only independent +variables are the input values \( x_n \).

    @@ -356,6 +369,9 @@ used as input to the activation functions. For each operation

  • 66
  • 67
  • 68
  • +
  • 69
  • +
  • 70
  • +
  • 71
  • »
  • diff --git a/doc/pub/week40/html/._week40-bs065.html b/doc/pub/week40/html/._week40-bs065.html index 9d5c14000..a1e59d665 100644 --- a/doc/pub/week40/html/._week40-bs065.html +++ b/doc/pub/week40/html/._week40-bs065.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -319,20 +334,32 @@ MathJax.Hub.Config({

     

     

     

    -

    Activation functions

    +

    Mathematical model

    -

    A property that characterizes a neural network, other than its -connectivity, is the choice of activation function(s). As described -in, the following restrictions are imposed on an activation function -for a FFNN to fulfill the universal approximation theorem +

    This confirms that an MLP, despite its quite convoluted mathematical +form, is nothing more than an analytic function, specifically a +mapping of real-valued vectors \( \hat{x} \in \mathbb{R}^n \rightarrow +\hat{y} \in \mathbb{R}^m \). +

    + +

    Furthermore, the flexibility and universality of an MLP can be +illustrated by realizing that the expression is essentially a nested +sum of scaled activation functions of the form +

    + +$$ +\begin{equation} + f(x) = c_1 f(c_2 x + c_3) + c_4 +\tag{15} +\end{equation} +$$ + +

    where the parameters \( c_i \) are weights and biases. By adjusting these +parameters, the activation functions can be shifted up and down or +left and right, change slope or be rescaled which is the key to the +flexibility of a neural network.

    -
      -
    • Non-constant
    • -
    • Bounded
    • -
    • Monotonically-increasing
    • -
    • Continuous
    • -

      @@ -350,6 +377,9 @@ for a FFNN to fulfill the universal approximation theorem
    • 66
    • 67
    • 68
    • +
    • 69
    • +
    • 70
    • +
    • 71
    • »
    diff --git a/doc/pub/week40/html/week40-bs.html b/doc/pub/week40/html/week40-bs.html index 5bc847095..70578a21e 100644 --- a/doc/pub/week40/html/week40-bs.html +++ b/doc/pub/week40/html/week40-bs.html @@ -37,6 +37,18 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d
  • Plans for week 40
  • -
  • Summary from last week, using gradient descent methods, limitations
  • -
  • Overview video on Stochastic Gradient Descent
  • -
  • Batches and mini-batches
  • -
  • Stochastic Gradient Descent (SGD)
  • -
  • Stochastic Gradient Descent
  • -
  • Computation of gradients
  • -
  • SGD example
  • -
  • The gradient step
  • -
  • Simple example code
  • -
  • When do we stop?
  • -
  • Slightly different approach
  • -
  • Time decay rate
  • -
  • Code with a Number of Minibatches which varies
  • -
  • Replace or not
  • -
  • Momentum based GD
  • -
  • More on momentum based approaches
  • -
  • Momentum parameter
  • -
  • Second moment of the gradient
  • -
  • RMS prop
  • -
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • -
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • -
  • Practical tips
  • -
  • Automatic differentiation
  • -
  • Using autograd
  • -
  • Autograd with more complicated functions
  • -
  • More complicated functions using the elements of their arguments directly
  • -
  • Functions using mathematical functions from Numpy
  • -
  • More autograd
  • -
  • And with loops
  • -
  • Using recursion
  • -
  • Unsupported functions
  • -
  • The syntax a.dot(b) when finding the dot product
  • -
  • Recommended to avoid
  • -
  • Using Autograd with OLS
  • -
  • Same code but now with momentum gradient descent
  • -
  • But noen of these can compete with Newton's method
  • -
  • Including Stochastic Gradient Descent with Autograd
  • -
  • Same code but now with momentum gradient descent
  • -
  • Similar (second order function now) problem but now with AdaGrad
  • -
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • -
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • -
  • And Logistic Regression
  • -
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • -
  • Introduction to Neural networks
  • -
  • Artificial neurons
  • -
  • Neural network types
  • -
  • Feed-forward neural networks
  • -
  • Convolutional Neural Network
  • -
  • Recurrent neural networks
  • -
  • Other types of networks
  • -
  • Multilayer perceptrons
  • -
  • Why multilayer perceptrons?
  • -
  • Illustration of a single perceptron model and a multi-perceptron model
  • -
  • Examples of XOR, OR and AND gates
  • -
  • Does Logistic Regression do a better Job?
  • -
  • Adding Neural Networks
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  • Mathematical model
  • -
  •    Matrix-vector notation
  • -
  •    Matrix-vector notation and activation
  • -
  •    Activation functions
  • -
  •    Activation functions, Logistic and Hyperbolic ones
  • -
  •    Relevance
  • +
  • Lecture Monday September 30, 2024
  • +
  • Suggested readings and videos
  • +
  • Lab sessions Tuesday and Wednesday
  • +
  • Summary from last week, using gradient descent methods, limitations
  • +
  • Overview video on Stochastic Gradient Descent
  • +
  • Batches and mini-batches
  • +
  • Stochastic Gradient Descent (SGD)
  • +
  • Stochastic Gradient Descent
  • +
  • Computation of gradients
  • +
  • SGD example
  • +
  • The gradient step
  • +
  • Simple example code
  • +
  • When do we stop?
  • +
  • Slightly different approach
  • +
  • Time decay rate
  • +
  • Code with a Number of Minibatches which varies
  • +
  • Replace or not
  • +
  • Momentum based GD
  • +
  • More on momentum based approaches
  • +
  • Momentum parameter
  • +
  • Second moment of the gradient
  • +
  • RMS prop
  • +
  • "ADAM optimizer":"https://arxiv.org/abs/1412.6980"
  • +
  • Algorithms and codes for Adagrad, RMSprop and Adam
  • +
  • Practical tips
  • +
  • Automatic differentiation
  • +
  • Using autograd
  • +
  • Autograd with more complicated functions
  • +
  • More complicated functions using the elements of their arguments directly
  • +
  • Functions using mathematical functions from Numpy
  • +
  • More autograd
  • +
  • And with loops
  • +
  • Using recursion
  • +
  • Unsupported functions
  • +
  • The syntax a.dot(b) when finding the dot product
  • +
  • Recommended to avoid
  • +
  • Using Autograd with OLS
  • +
  • Same code but now with momentum gradient descent
  • +
  • But noen of these can compete with Newton's method
  • +
  • Including Stochastic Gradient Descent with Autograd
  • +
  • Same code but now with momentum gradient descent
  • +
  • Similar (second order function now) problem but now with AdaGrad
  • +
  • RMSprop for adaptive learning rate with Stochastic Gradient Descent
  • +
  • And finally "ADAM":"https://arxiv.org/pdf/1412.6980.pdf"
  • +
  • And Logistic Regression
  • +
  • Introducing "JAX":"https://jax.readthedocs.io/en/latest/"
  • +
  • Introduction to Neural networks
  • +
  • Artificial neurons
  • +
  • Neural network types
  • +
  • Feed-forward neural networks
  • +
  • Convolutional Neural Network
  • +
  • Recurrent neural networks
  • +
  • Other types of networks
  • +
  • Multilayer perceptrons
  • +
  • Why multilayer perceptrons?
  • +
  • Illustration of a single perceptron model and a multi-perceptron model
  • +
  • Examples of XOR, OR and AND gates
  • +
  • Does Logistic Regression do a better Job?
  • +
  • Adding Neural Networks
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  • Mathematical model
  • +
  •    Matrix-vector notation
  • +
  •    Matrix-vector notation and activation
  • +
  •    Activation functions
  • +
  •    Activation functions, Logistic and Hyperbolic ones
  • +
  •    Relevance
  • @@ -337,7 +352,7 @@ MathJax.Hub.Config({
    -

    October 2-6, 2023

    +

    September 30-October 4, 2024


    @@ -362,7 +377,7 @@ MathJax.Hub.Config({
  • 9
  • 10
  • ...
  • -
  • 68
  • +
  • 71
  • »
  • @@ -376,7 +391,7 @@ MathJax.Hub.Config({ -->
    - © 1999-2023, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license + © 1999-2024, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license
    diff --git a/doc/pub/week40/html/week40-reveal.html b/doc/pub/week40/html/week40-reveal.html index f2e6c76fe..6911c6ca8 100644 --- a/doc/pub/week40/html/week40-reveal.html +++ b/doc/pub/week40/html/week40-reveal.html @@ -184,19 +184,60 @@ MathJax.Hub.Config({
    -

    October 2-6, 2023

    +

    September 30-October 4, 2024


    - © 1999-2023, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license + © 1999-2024, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license

    Plans for week 40

    +
    +
    +

    Lecture Monday September 30, 2024

    +
    + +

    +

      +

    1. Stochastic Gradient descent with examples and automatic differentiation
    2. +

    3. If we get time, we start with the basics of Neural Networks, setting up the basic steps, from the simple perceptron model to the multi-layer perceptron model + +
    4. +
    +
    +
    + +
    +

    Suggested readings and videos

    +
    +Readings and Videos: +

    +

      + +

    1. The lecture notes for week 40 (these notes)
    2. + +

    3. For a good discussion on gradient methods, we would like to recommend Goodfellow et al section 4.3-4.5 and sections 8.3-8.6. We will come back to the latter chapter in our discussion of Neural networks as well.
    4. + +

    5. For neural networks we recommend Goodfellow et al chapter 6 and Raschka et al pages 48-60
    6. + +

    7. Video on gradient descent at https://www.youtube.com/watch?v=sDv4f4s2SB8
    8. + +

    9. Video on stochastic gradient descent at https://www.youtube.com/watch?v=vMh0zPT0tLI
    10. + +

    11. Neural Networks demystified at https://www.youtube.com/watch?v=bxe2T-V8XRs&list=PLiaHhY2iBX9hdHaRr6b7XevZtgZRa1PoU&ab_channel=WelchLabs
    12. + +

    13. Building Neural Networks from scratch at URL:https://www.youtube.com/watch?v=Wo5dMEP_BbI&list=PLQVvvaa0QuDcjD5BAw2DxE6OF2tius3V3&ab_channel=sentdex"
    14. +
    +
    +
    + +
    +

    Lab sessions Tuesday and Wednesday

    Material for the active learning sessions on Tuesday and Wednesday

    @@ -206,48 +247,11 @@ MathJax.Hub.Config({

  • No weekly exercises for week 40, project work only
  • -

  • Video on how to write scientific reports recorded during one of the lab sessions
  • +

  • Video on how to write scientific reports recorded during one of the lab sessions at https://youtu.be/tVW1ZDmZnwM
  • A general guideline can be found at https://github.com/CompPhysics/MachineLearning/blob/master/doc/Projects/EvaluationGrading/EvaluationForm.md.
  • - - -
    -Material for the lecture on Thursday October 5, 2023 -

    -

    -
    @@ -793,7 +797,7 @@ the steep computational price of calculating or approximating Hessians.

    -

    Recently, a number of methods have been introduced that accomplish +

    During the last decade a number of methods have been introduced that accomplish this by tracking not only the gradient, but also the second moment of the gradient. These methods include AdaGrad, AdaDelta, Root Mean Squared Propagation (RMS-Prop), and ADAM. diff --git a/doc/pub/week40/html/week40-solarized.html b/doc/pub/week40/html/week40-solarized.html index 0b0ee1ee8..5ddadc456 100644 --- a/doc/pub/week40/html/week40-solarized.html +++ b/doc/pub/week40/html/week40-solarized.html @@ -64,6 +64,18 @@ div.toc p,a {










    Plans for week 40

    +









    +

    Lecture Monday September 30, 2024

    +
    + +

    +

      +
    1. Stochastic Gradient descent with examples and automatic differentiation
    2. +
    3. If we get time, we start with the basics of Neural Networks, setting up the basic steps, from the simple perceptron model to the multi-layer perceptron model + +
    4. +
    +
    + + +









    +

    Suggested readings and videos

    +
    +Readings and Videos: +

    +

      +
    1. The lecture notes for week 40 (these notes)
    2. +
    3. For a good discussion on gradient methods, we would like to recommend Goodfellow et al section 4.3-4.5 and sections 8.3-8.6. We will come back to the latter chapter in our discussion of Neural networks as well.
    4. +
    5. For neural networks we recommend Goodfellow et al chapter 6 and Raschka et al pages 48-60
    6. +
    7. Video on gradient descent at https://www.youtube.com/watch?v=sDv4f4s2SB8
    8. +
    9. Video on stochastic gradient descent at https://www.youtube.com/watch?v=vMh0zPT0tLI
    10. +
    11. Neural Networks demystified at https://www.youtube.com/watch?v=bxe2T-V8XRs&list=PLiaHhY2iBX9hdHaRr6b7XevZtgZRa1PoU&ab_channel=WelchLabs
    12. +
    13. Building Neural Networks from scratch at URL:https://www.youtube.com/watch?v=Wo5dMEP_BbI&list=PLQVvvaa0QuDcjD5BAw2DxE6OF2tius3V3&ab_channel=sentdex"
    14. +
    +
    + + +









    +

    Lab sessions Tuesday and Wednesday

    Material for the active learning sessions on Tuesday and Wednesday

    -
    -Material for the lecture on Thursday October 5, 2023 -

    -

    -
    - -









    Summary from last week, using gradient descent methods, limitations

    @@ -819,7 +841,7 @@ the steep computational price of calculating or approximating Hessians.

    -

    Recently, a number of methods have been introduced that accomplish +

    During the last decade a number of methods have been introduced that accomplish this by tracking not only the gradient, but also the second moment of the gradient. These methods include AdaGrad, AdaDelta, Root Mean Squared Propagation (RMS-Prop), and ADAM. @@ -3025,7 +3047,7 @@ plt.show()

    - © 1999-2023, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license + © 1999-2024, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license
    diff --git a/doc/pub/week40/html/week40.html b/doc/pub/week40/html/week40.html index 1daf0eb0f..2f3ca087a 100644 --- a/doc/pub/week40/html/week40.html +++ b/doc/pub/week40/html/week40.html @@ -141,6 +141,18 @@ div.toc p,a {










    Plans for week 40

    +









    +

    Lecture Monday September 30, 2024

    +
    + +

    +

      +
    1. Stochastic Gradient descent with examples and automatic differentiation
    2. +
    3. If we get time, we start with the basics of Neural Networks, setting up the basic steps, from the simple perceptron model to the multi-layer perceptron model + +
    4. +
    +
    + + +









    +

    Suggested readings and videos

    +
    +Readings and Videos: +

    +

      +
    1. The lecture notes for week 40 (these notes)
    2. +
    3. For a good discussion on gradient methods, we would like to recommend Goodfellow et al section 4.3-4.5 and sections 8.3-8.6. We will come back to the latter chapter in our discussion of Neural networks as well.
    4. +
    5. For neural networks we recommend Goodfellow et al chapter 6 and Raschka et al pages 48-60
    6. +
    7. Video on gradient descent at https://www.youtube.com/watch?v=sDv4f4s2SB8
    8. +
    9. Video on stochastic gradient descent at https://www.youtube.com/watch?v=vMh0zPT0tLI
    10. +
    11. Neural Networks demystified at https://www.youtube.com/watch?v=bxe2T-V8XRs&list=PLiaHhY2iBX9hdHaRr6b7XevZtgZRa1PoU&ab_channel=WelchLabs
    12. +
    13. Building Neural Networks from scratch at URL:https://www.youtube.com/watch?v=Wo5dMEP_BbI&list=PLQVvvaa0QuDcjD5BAw2DxE6OF2tius3V3&ab_channel=sentdex"
    14. +
    +
    + + +









    +

    Lab sessions Tuesday and Wednesday

    Material for the active learning sessions on Tuesday and Wednesday

    -
    -Material for the lecture on Thursday October 5, 2023 -

    -

    -
    - -









    Summary from last week, using gradient descent methods, limitations

    @@ -896,7 +918,7 @@ the steep computational price of calculating or approximating Hessians.

    -

    Recently, a number of methods have been introduced that accomplish +

    During the last decade a number of methods have been introduced that accomplish this by tracking not only the gradient, but also the second moment of the gradient. These methods include AdaGrad, AdaDelta, Root Mean Squared Propagation (RMS-Prop), and ADAM. @@ -3102,7 +3124,7 @@ plt.show()

    - © 1999-2023, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license + © 1999-2024, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license
    diff --git a/doc/pub/week40/ipynb/ipynb-week40-src.tar.gz b/doc/pub/week40/ipynb/ipynb-week40-src.tar.gz index c11c12f249236d783a42c49a56bbab9d816e7bda..aadf9707d2bd7022e4a964fb318a78e1441cf734 100644 GIT binary patch delta 19 acmbQ%!!)gji9^1dgW>IppJ^L8_}TzQRtH1? delta 19 acmbQ%!!)gji9^1dgQ4t{T\n", + "" + ] + }, + { + "cell_type": "markdown", + "id": "f9e79497", + "metadata": { + "editable": true + }, + "source": [ + "## Suggested readings and videos\n", + "**Readings and Videos:**\n", + "\n", + "1. The lecture notes for week 40 (these notes)\n", + "\n", + "2. For a good discussion on gradient methods, we would like to recommend Goodfellow et al section 4.3-4.5 and sections 8.3-8.6. We will come back to the latter chapter in our discussion of Neural networks as well.\n", + "\n", + "3. For neural networks we recommend Goodfellow et al chapter 6 and Raschka et al pages 48-60\n", + "\n", + "4. Video on gradient descent at \n", + "\n", + "5. Video on stochastic gradient descent at \n", + "\n", + "6. Neural Networks demystified at \n", + "\n", + "7. Building Neural Networks from scratch at URL:https://www.youtube.com/watch?v=Wo5dMEP_BbI&list=PLQVvvaa0QuDcjD5BAw2DxE6OF2tius3V3&ab_channel=sentdex\"" + ] + }, + { + "cell_type": "markdown", + "id": "2a48a209", + "metadata": { + "editable": true + }, + "source": [ + "## Lab sessions Tuesday and Wednesday\n", "**Material for the active learning sessions on Tuesday and Wednesday.**\n", "\n", " * Work on project 1 and discussions on how to structure your report\n", "\n", " * No weekly exercises for week 40, project work only\n", "\n", - " * [Video on how to write scientific reports recorded during one of the lab sessions](https://youtu.be/tVW1ZDmZnwM)\n", + " * Video on how to write scientific reports recorded during one of the lab sessions at \n", "\n", - " * A general guideline can be found at .\n", - "\n", - " \n", - "\n", - "**Material for the lecture on Thursday October 5, 2023.**\n", - "\n", - " * Stochastic Gradient descent with examples and automatic differentiation\n", - "\n", - " * Neural Networks, setting up the basic steps, from the simple perceptron model to the multi-layer perceptron model.\n", - "\n", - " * [Video of lecture](https://youtu.be/75pr3hKY20U)\n", - "\n", - " * \"Whiteboard notes at \n", - "\n", - " * Readings and Videos:\n", - "\n", - " * These lecture notes\n", - "\n", - " * For a good discussion on gradient methods, we would like to recommend Goodfellow et al section 4.3-4.5 and sections 8.3-8.6. We will come back to the latter chapter in our discussion of Neural networks as well.\n", - "\n", - " * [Aurelien Geron's chapter 4 on stochastic gradient descent](https://github.com/CompPhysics/MachineLearning/blob/master/doc/Textbooks/TensorflowML.pdf)\n", - "\n", - " * For neural networks we recommend Goodfellow et al chapter 6.\n", - "\n", - " * [Video on gradient descent](https://www.youtube.com/watch?v=sDv4f4s2SB8)\n", - "\n", - " * [Video on stochastic gradient descent](https://www.youtube.com/watch?v=vMh0zPT0tLI)\n", - "\n", - " * [Neural Networks demystified](https://www.youtube.com/watch?v=bxe2T-V8XRs&list=PLiaHhY2iBX9hdHaRr6b7XevZtgZRa1PoU&ab_channel=WelchLabs)\n", - "\n", - " * [Building Neural Networks from scratch](https://www.youtube.com/watch?v=Wo5dMEP_BbI&list=PLQVvvaa0QuDcjD5BAw2DxE6OF2tius3V3&ab_channel=sentdex)" + " * A general guideline can be found at ." ] }, { "cell_type": "markdown", - "id": "ef0a6cd1", + "id": "9438015b", "metadata": { "editable": true }, @@ -99,7 +118,7 @@ }, { "cell_type": "markdown", - "id": "be641f47", + "id": "7effc5b6", "metadata": { "editable": true }, @@ -111,7 +130,7 @@ }, { "cell_type": "markdown", - "id": "f4b83555", + "id": "3ab97704", "metadata": { "editable": true }, @@ -132,7 +151,7 @@ }, { "cell_type": "markdown", - "id": "92eb60ea", + "id": "98025b12", "metadata": { "editable": true }, @@ -164,7 +183,7 @@ }, { "cell_type": "markdown", - "id": "ff93cc6a", + "id": "9d22c1c2", "metadata": { "editable": true }, @@ -181,7 +200,7 @@ }, { "cell_type": "markdown", - "id": "a9a6056b", + "id": "25b41726", "metadata": { "editable": true }, @@ -194,7 +213,7 @@ }, { "cell_type": "markdown", - "id": "34bdb509", + "id": "3f6df491", "metadata": { "editable": true }, @@ -207,7 +226,7 @@ }, { "cell_type": "markdown", - "id": "21dd7d3e", + "id": "456219de", "metadata": { "editable": true }, @@ -220,7 +239,7 @@ }, { "cell_type": "markdown", - "id": "50d2de84", + "id": "4502050f", "metadata": { "editable": true }, @@ -234,7 +253,7 @@ }, { "cell_type": "markdown", - "id": "35c3d710", + "id": "c3d9ce8c", "metadata": { "editable": true }, @@ -256,7 +275,7 @@ }, { "cell_type": "markdown", - "id": "edae861f", + "id": "c120a02c", "metadata": { "editable": true }, @@ -271,7 +290,7 @@ }, { "cell_type": "markdown", - "id": "d3ae1891", + "id": "becac233", "metadata": { "editable": true }, @@ -283,7 +302,7 @@ }, { "cell_type": "markdown", - "id": "03e07c8b", + "id": "4868b046", "metadata": { "editable": true }, @@ -296,7 +315,7 @@ }, { "cell_type": "markdown", - "id": "db13df10", + "id": "80031167", "metadata": { "editable": true }, @@ -310,7 +329,7 @@ }, { "cell_type": "markdown", - "id": "10942435", + "id": "1bf35f7e", "metadata": { "editable": true }, @@ -321,7 +340,7 @@ { "cell_type": "code", "execution_count": 1, - "id": "83ee2bfc", + "id": "5daa08f8", "metadata": { "collapsed": false, "editable": true @@ -346,7 +365,7 @@ }, { "cell_type": "markdown", - "id": "bc1a309e", + "id": "4277c834", "metadata": { "editable": true }, @@ -362,7 +381,7 @@ }, { "cell_type": "markdown", - "id": "20ebc0b1", + "id": "924c0b23", "metadata": { "editable": true }, @@ -383,7 +402,7 @@ }, { "cell_type": "markdown", - "id": "33b59a14", + "id": "3dd0da01", "metadata": { "editable": true }, @@ -403,7 +422,7 @@ }, { "cell_type": "markdown", - "id": "55aef65a", + "id": "d59f421e", "metadata": { "editable": true }, @@ -422,7 +441,7 @@ { "cell_type": "code", "execution_count": 2, - "id": "c0b3004c", + "id": "e459d66b", "metadata": { "collapsed": false, "editable": true @@ -457,7 +476,7 @@ }, { "cell_type": "markdown", - "id": "3db5b62f", + "id": "4c24e703", "metadata": { "editable": true }, @@ -470,7 +489,7 @@ { "cell_type": "code", "execution_count": 3, - "id": "96488d3f", + "id": "c1eb4f7e", "metadata": { "collapsed": false, "editable": true @@ -549,7 +568,7 @@ }, { "cell_type": "markdown", - "id": "228b92a4", + "id": "1b91fb30", "metadata": { "editable": true }, @@ -564,7 +583,7 @@ }, { "cell_type": "markdown", - "id": "4465e353", + "id": "ab479a58", "metadata": { "editable": true }, @@ -579,7 +598,7 @@ }, { "cell_type": "markdown", - "id": "0463b6b5", + "id": "3b0eec5e", "metadata": { "editable": true }, @@ -591,7 +610,7 @@ }, { "cell_type": "markdown", - "id": "6e9ec0b0", + "id": "7a184369", "metadata": { "editable": true }, @@ -609,7 +628,7 @@ }, { "cell_type": "markdown", - "id": "75dfe3e3", + "id": "03a0ba9d", "metadata": { "editable": true }, @@ -628,7 +647,7 @@ }, { "cell_type": "markdown", - "id": "f290803e", + "id": "76301f54", "metadata": { "editable": true }, @@ -640,7 +659,7 @@ }, { "cell_type": "markdown", - "id": "46975656", + "id": "225f6ab3", "metadata": { "editable": true }, @@ -650,7 +669,7 @@ }, { "cell_type": "markdown", - "id": "42a25740", + "id": "fef1e968", "metadata": { "editable": true }, @@ -666,7 +685,7 @@ }, { "cell_type": "markdown", - "id": "08ce02e9", + "id": "950bb7cd", "metadata": { "editable": true }, @@ -678,7 +697,7 @@ }, { "cell_type": "markdown", - "id": "24772dd9", + "id": "b9e6646d", "metadata": { "editable": true }, @@ -688,7 +707,7 @@ }, { "cell_type": "markdown", - "id": "6a7e21b1", + "id": "26eefa05", "metadata": { "editable": true }, @@ -700,7 +719,7 @@ }, { "cell_type": "markdown", - "id": "026b037e", + "id": "f9747238", "metadata": { "editable": true }, @@ -710,7 +729,7 @@ }, { "cell_type": "markdown", - "id": "df89140f", + "id": "800512f6", "metadata": { "editable": true }, @@ -722,7 +741,7 @@ }, { "cell_type": "markdown", - "id": "b250b334", + "id": "b1f742be", "metadata": { "editable": true }, @@ -738,7 +757,7 @@ }, { "cell_type": "markdown", - "id": "f3e99e54", + "id": "8e6fea67", "metadata": { "editable": true }, @@ -750,7 +769,7 @@ }, { "cell_type": "markdown", - "id": "4d250128", + "id": "0c3e1b5a", "metadata": { "editable": true }, @@ -783,7 +802,7 @@ }, { "cell_type": "markdown", - "id": "074867f5", + "id": "3efff2e2", "metadata": { "editable": true }, @@ -795,7 +814,7 @@ }, { "cell_type": "markdown", - "id": "577c3314", + "id": "18764ecd", "metadata": { "editable": true }, @@ -813,7 +832,7 @@ }, { "cell_type": "markdown", - "id": "87d3fffb", + "id": "4e900ba9", "metadata": { "editable": true }, @@ -823,7 +842,7 @@ }, { "cell_type": "markdown", - "id": "5e484c2a", + "id": "0c4bcbe4", "metadata": { "editable": true }, @@ -846,7 +865,7 @@ "the steep computational price of calculating or approximating\n", "Hessians.\n", "\n", - "Recently, a number of methods have been introduced that accomplish\n", + "During the last decade a number of methods have been introduced that accomplish\n", "this by tracking not only the gradient, but also the second moment of\n", "the gradient. These methods include AdaGrad, AdaDelta, Root Mean Squared Propagation (RMS-Prop), and\n", "[ADAM](https://arxiv.org/abs/1412.6980)." @@ -854,7 +873,7 @@ }, { "cell_type": "markdown", - "id": "812f0e0e", + "id": "91e23db3", "metadata": { "editable": true }, @@ -869,7 +888,7 @@ }, { "cell_type": "markdown", - "id": "ffbb49b0", + "id": "c0b3c862", "metadata": { "editable": true }, @@ -887,7 +906,7 @@ }, { "cell_type": "markdown", - "id": "8471bf35", + "id": "587eb65d", "metadata": { "editable": true }, @@ -899,7 +918,7 @@ }, { "cell_type": "markdown", - "id": "ac7a25ba", + "id": "fc251cfa", "metadata": { "editable": true }, @@ -911,7 +930,7 @@ }, { "cell_type": "markdown", - "id": "170cec32", + "id": "5f940a4b", "metadata": { "editable": true }, @@ -929,7 +948,7 @@ }, { "cell_type": "markdown", - "id": "603a78e2", + "id": "b18e2cda", "metadata": { "editable": true }, @@ -958,7 +977,7 @@ }, { "cell_type": "markdown", - "id": "458870be", + "id": "ee0eea90", "metadata": { "editable": true }, @@ -976,7 +995,7 @@ }, { "cell_type": "markdown", - "id": "807e4dc9", + "id": "1210bffb", "metadata": { "editable": true }, @@ -988,7 +1007,7 @@ }, { "cell_type": "markdown", - "id": "62479ff0", + "id": "dc3c0521", "metadata": { "editable": true }, @@ -1000,7 +1019,7 @@ }, { "cell_type": "markdown", - "id": "eda102d5", + "id": "a9c7c5e6", "metadata": { "editable": true }, @@ -1012,7 +1031,7 @@ }, { "cell_type": "markdown", - "id": "07c5278d", + "id": "38fd26bd", "metadata": { "editable": true }, @@ -1024,7 +1043,7 @@ }, { "cell_type": "markdown", - "id": "cf18307f", + "id": "3cdc2058", "metadata": { "editable": true }, @@ -1036,7 +1055,7 @@ }, { "cell_type": "markdown", - "id": "2c4bac72", + "id": "00ecf3cc", "metadata": { "editable": true }, @@ -1053,7 +1072,7 @@ }, { "cell_type": "markdown", - "id": "a2670df6", + "id": "b7ec1ddf", "metadata": { "editable": true }, @@ -1072,7 +1091,7 @@ }, { "cell_type": "markdown", - "id": "165cf946", + "id": "9758c4f9", "metadata": { "editable": true }, @@ -1084,7 +1103,7 @@ }, { "cell_type": "markdown", - "id": "ad4252c9", + "id": "dcfce298", "metadata": { "editable": true }, @@ -1098,7 +1117,7 @@ }, { "cell_type": "markdown", - "id": "8ab76031", + "id": "e8b19969", "metadata": { "editable": true }, @@ -1118,7 +1137,7 @@ }, { "cell_type": "markdown", - "id": "4d78ebb5", + "id": "445f1f26", "metadata": { "editable": true }, @@ -1156,7 +1175,7 @@ }, { "cell_type": "markdown", - "id": "1c9e695b", + "id": "e5edb15a", "metadata": { "editable": true }, @@ -1168,7 +1187,7 @@ }, { "cell_type": "markdown", - "id": "7e26ad18", + "id": "f7563c65", "metadata": { "editable": true }, @@ -1178,7 +1197,7 @@ }, { "cell_type": "markdown", - "id": "92894a23", + "id": "7f6d1beb", "metadata": { "editable": true }, @@ -1190,7 +1209,7 @@ }, { "cell_type": "markdown", - "id": "c8de4271", + "id": "6c8a48fb", "metadata": { "editable": true }, @@ -1201,7 +1220,7 @@ { "cell_type": "code", "execution_count": 4, - "id": "7ea01e86", + "id": "d52663b9", "metadata": { "collapsed": false, "editable": true @@ -1246,7 +1265,7 @@ }, { "cell_type": "markdown", - "id": "ae059e6d", + "id": "a1fe986b", "metadata": { "editable": true }, @@ -1263,7 +1282,7 @@ { "cell_type": "code", "execution_count": 5, - "id": "09c98936", + "id": "f31996b4", "metadata": { "collapsed": false, "editable": true @@ -1291,7 +1310,7 @@ }, { "cell_type": "markdown", - "id": "cca3b36d", + "id": "3cf6c84b", "metadata": { "editable": true }, @@ -1306,7 +1325,7 @@ { "cell_type": "code", "execution_count": 6, - "id": "30136aad", + "id": "0ef4e00e", "metadata": { "collapsed": false, "editable": true @@ -1350,7 +1369,7 @@ }, { "cell_type": "markdown", - "id": "d1302515", + "id": "74ff4efc", "metadata": { "editable": true }, @@ -1360,7 +1379,7 @@ }, { "cell_type": "markdown", - "id": "5d5e5166", + "id": "2c3c722f", "metadata": { "editable": true }, @@ -1371,7 +1390,7 @@ { "cell_type": "code", "execution_count": 7, - "id": "3659886d", + "id": "88926080", "metadata": { "collapsed": false, "editable": true @@ -1399,7 +1418,7 @@ }, { "cell_type": "markdown", - "id": "81f95eed", + "id": "b6e58b5f", "metadata": { "editable": true }, @@ -1414,7 +1433,7 @@ }, { "cell_type": "markdown", - "id": "8020598c", + "id": "af5d0ed7", "metadata": { "editable": true }, @@ -1425,7 +1444,7 @@ { "cell_type": "code", "execution_count": 8, - "id": "a13dd118", + "id": "c85bb478", "metadata": { "collapsed": false, "editable": true @@ -1453,7 +1472,7 @@ }, { "cell_type": "markdown", - "id": "7cd1b75e", + "id": "9322ae7d", "metadata": { "editable": true }, @@ -1464,7 +1483,7 @@ { "cell_type": "code", "execution_count": 9, - "id": "3118ffa4", + "id": "e7772208", "metadata": { "collapsed": false, "editable": true @@ -1489,7 +1508,7 @@ }, { "cell_type": "markdown", - "id": "b088c610", + "id": "db7c01b0", "metadata": { "editable": true }, @@ -1500,7 +1519,7 @@ { "cell_type": "code", "execution_count": 10, - "id": "3af27b63", + "id": "f9cbb48d", "metadata": { "collapsed": false, "editable": true @@ -1536,7 +1555,7 @@ { "cell_type": "code", "execution_count": 11, - "id": "0b8db314", + "id": "8ad91d6c", "metadata": { "collapsed": false, "editable": true @@ -1556,7 +1575,7 @@ }, { "cell_type": "markdown", - "id": "4e49e39a", + "id": "5b074c31", "metadata": { "editable": true }, @@ -1567,7 +1586,7 @@ { "cell_type": "code", "execution_count": 12, - "id": "67954e7f", + "id": "50139508", "metadata": { "collapsed": false, "editable": true @@ -1605,7 +1624,7 @@ }, { "cell_type": "markdown", - "id": "af5d8d0f", + "id": "ea244d9a", "metadata": { "editable": true }, @@ -1615,7 +1634,7 @@ }, { "cell_type": "markdown", - "id": "dbdbf611", + "id": "4f564a15", "metadata": { "editable": true }, @@ -1629,7 +1648,7 @@ { "cell_type": "code", "execution_count": 13, - "id": "f5475c2d", + "id": "e192cad7", "metadata": { "collapsed": false, "editable": true @@ -1651,7 +1670,7 @@ }, { "cell_type": "markdown", - "id": "b330a678", + "id": "0929ecc1", "metadata": { "editable": true }, @@ -1661,7 +1680,7 @@ }, { "cell_type": "markdown", - "id": "d71f3cee", + "id": "d69a4e44", "metadata": { "editable": true }, @@ -1672,7 +1691,7 @@ { "cell_type": "code", "execution_count": 14, - "id": "bf466879", + "id": "5f67b908", "metadata": { "collapsed": false, "editable": true @@ -1694,7 +1713,7 @@ }, { "cell_type": "markdown", - "id": "da58ce6d", + "id": "66bcfa89", "metadata": { "editable": true }, @@ -1707,7 +1726,7 @@ { "cell_type": "code", "execution_count": 15, - "id": "1d82b15c", + "id": "d4bca4d5", "metadata": { "collapsed": false, "editable": true @@ -1732,7 +1751,7 @@ }, { "cell_type": "markdown", - "id": "6f888ca5", + "id": "77fe70ec", "metadata": { "editable": true }, @@ -1744,7 +1763,7 @@ { "cell_type": "code", "execution_count": 16, - "id": "d2049347", + "id": "8a861e1b", "metadata": { "collapsed": false, "editable": true @@ -1759,7 +1778,7 @@ }, { "cell_type": "markdown", - "id": "23df479a", + "id": "d2408e10", "metadata": { "editable": true }, @@ -1774,7 +1793,7 @@ { "cell_type": "code", "execution_count": 17, - "id": "b00d07b0", + "id": "3cf3469a", "metadata": { "collapsed": false, "editable": true @@ -1834,7 +1853,7 @@ }, { "cell_type": "markdown", - "id": "7c8ebe00", + "id": "6c79f605", "metadata": { "editable": true }, @@ -1845,7 +1864,7 @@ { "cell_type": "code", "execution_count": 18, - "id": "589722bc", + "id": "abfa926a", "metadata": { "collapsed": false, "editable": true @@ -1909,7 +1928,7 @@ }, { "cell_type": "markdown", - "id": "f89c3bea", + "id": "82247c7a", "metadata": { "editable": true }, @@ -1920,7 +1939,7 @@ { "cell_type": "code", "execution_count": 19, - "id": "0016cfef", + "id": "aeb8ebda", "metadata": { "collapsed": false, "editable": true @@ -1969,7 +1988,7 @@ }, { "cell_type": "markdown", - "id": "40c08697", + "id": "d76b6beb", "metadata": { "editable": true }, @@ -1981,7 +2000,7 @@ { "cell_type": "code", "execution_count": 20, - "id": "0ac46269", + "id": "de526571", "metadata": { "collapsed": false, "editable": true @@ -2065,7 +2084,7 @@ }, { "cell_type": "markdown", - "id": "5b54e40c", + "id": "0d89e0bf", "metadata": { "editable": true }, @@ -2076,7 +2095,7 @@ { "cell_type": "code", "execution_count": 21, - "id": "9816ac5a", + "id": "620345fc", "metadata": { "collapsed": false, "editable": true @@ -2154,7 +2173,7 @@ }, { "cell_type": "markdown", - "id": "6575a5e2", + "id": "caf732a2", "metadata": { "editable": true }, @@ -2165,7 +2184,7 @@ { "cell_type": "code", "execution_count": 22, - "id": "6189713a", + "id": "47049d27", "metadata": { "collapsed": false, "editable": true @@ -2224,7 +2243,7 @@ }, { "cell_type": "markdown", - "id": "08932a41", + "id": "189ca4e7", "metadata": { "editable": true }, @@ -2234,7 +2253,7 @@ }, { "cell_type": "markdown", - "id": "e17dc12b", + "id": "82019114", "metadata": { "editable": true }, @@ -2245,7 +2264,7 @@ { "cell_type": "code", "execution_count": 23, - "id": "af328676", + "id": "c97f1e0c", "metadata": { "collapsed": false, "editable": true @@ -2310,7 +2329,7 @@ }, { "cell_type": "markdown", - "id": "59ccc3be", + "id": "7cb1b5ac", "metadata": { "editable": true }, @@ -2321,7 +2340,7 @@ { "cell_type": "code", "execution_count": 24, - "id": "c080b34f", + "id": "6dca5827", "metadata": { "collapsed": false, "editable": true @@ -2391,7 +2410,7 @@ }, { "cell_type": "markdown", - "id": "9b32244d", + "id": "01339078", "metadata": { "editable": true }, @@ -2402,7 +2421,7 @@ { "cell_type": "code", "execution_count": 25, - "id": "5843a612", + "id": "08670401", "metadata": { "collapsed": false, "editable": true @@ -2446,7 +2465,7 @@ }, { "cell_type": "markdown", - "id": "a810cd04", + "id": "8161e830", "metadata": { "editable": true }, @@ -2465,7 +2484,7 @@ { "cell_type": "code", "execution_count": 26, - "id": "a83fb9f5", + "id": "a834eb14", "metadata": { "collapsed": false, "editable": true @@ -2485,7 +2504,7 @@ }, { "cell_type": "markdown", - "id": "f10d8b41", + "id": "2f919ac6", "metadata": { "editable": true }, @@ -2503,7 +2522,7 @@ }, { "cell_type": "markdown", - "id": "e7631ec5", + "id": "b36ded80", "metadata": { "editable": true }, @@ -2527,7 +2546,7 @@ }, { "cell_type": "markdown", - "id": "d7faf9c9", + "id": "dae3f4b1", "metadata": { "editable": true }, @@ -2545,7 +2564,7 @@ }, { "cell_type": "markdown", - "id": "49575f6e", + "id": "cb361311", "metadata": { "editable": true }, @@ -2585,7 +2604,7 @@ }, { "cell_type": "markdown", - "id": "ca7aa5dc", + "id": "b4213e00", "metadata": { "editable": true }, @@ -2614,7 +2633,7 @@ }, { "cell_type": "markdown", - "id": "37020198", + "id": "9bdfe95b", "metadata": { "editable": true }, @@ -2635,7 +2654,7 @@ }, { "cell_type": "markdown", - "id": "09a662fc", + "id": "fcb7a268", "metadata": { "editable": true }, @@ -2664,7 +2683,7 @@ }, { "cell_type": "markdown", - "id": "37634c37", + "id": "d764f349", "metadata": { "editable": true }, @@ -2685,7 +2704,7 @@ }, { "cell_type": "markdown", - "id": "69f9e942", + "id": "39e52765", "metadata": { "editable": true }, @@ -2706,7 +2725,7 @@ }, { "cell_type": "markdown", - "id": "e60029c8", + "id": "6686b16c", "metadata": { "editable": true }, @@ -2723,7 +2742,7 @@ }, { "cell_type": "markdown", - "id": "3e6a2e84", + "id": "e150cc65", "metadata": { "editable": true }, @@ -2744,7 +2763,7 @@ }, { "cell_type": "markdown", - "id": "eee2fd77", + "id": "b9c29186", "metadata": { "editable": true }, @@ -2760,7 +2779,7 @@ }, { "cell_type": "markdown", - "id": "d732156c", + "id": "5ef78c3a", "metadata": { "editable": true }, @@ -2777,7 +2796,7 @@ { "cell_type": "code", "execution_count": 27, - "id": "f3b86313", + "id": "34aff373", "metadata": { "collapsed": false, "editable": true @@ -2818,7 +2837,7 @@ }, { "cell_type": "markdown", - "id": "86fcb8b9", + "id": "83522b2b", "metadata": { "editable": true }, @@ -2828,7 +2847,7 @@ }, { "cell_type": "markdown", - "id": "8f74c3b9", + "id": "7bf688ce", "metadata": { "editable": true }, @@ -2839,7 +2858,7 @@ { "cell_type": "code", "execution_count": 28, - "id": "0495d647", + "id": "97c28d21", "metadata": { "collapsed": false, "editable": true @@ -2899,7 +2918,7 @@ }, { "cell_type": "markdown", - "id": "afd47b0f", + "id": "5ec8385d", "metadata": { "editable": true }, @@ -2909,7 +2928,7 @@ }, { "cell_type": "markdown", - "id": "e1704d2c", + "id": "ddc2f5f6", "metadata": { "editable": true }, @@ -2920,7 +2939,7 @@ { "cell_type": "code", "execution_count": 29, - "id": "b344cc20", + "id": "94eba161", "metadata": { "collapsed": false, "editable": true @@ -2940,7 +2959,7 @@ }, { "cell_type": "markdown", - "id": "e39d6c65", + "id": "ddeb95ad", "metadata": { "editable": true }, @@ -2952,7 +2971,7 @@ }, { "cell_type": "markdown", - "id": "3d5bd705", + "id": "c9167ccd", "metadata": { "editable": true }, @@ -2964,7 +2983,7 @@ }, { "cell_type": "markdown", - "id": "9f426867", + "id": "be152c30", "metadata": { "editable": true }, @@ -2979,7 +2998,7 @@ }, { "cell_type": "markdown", - "id": "f1fc52b4", + "id": "2efc9bc5", "metadata": { "editable": true }, @@ -2991,7 +3010,7 @@ }, { "cell_type": "markdown", - "id": "0a37423c", + "id": "0f9e0ee7", "metadata": { "editable": true }, @@ -3008,7 +3027,7 @@ }, { "cell_type": "markdown", - "id": "c2bcc357", + "id": "dd3fd6b0", "metadata": { "editable": true }, @@ -3023,7 +3042,7 @@ }, { "cell_type": "markdown", - "id": "f7bc2cce", + "id": "32158edc", "metadata": { "editable": true }, @@ -3041,7 +3060,7 @@ }, { "cell_type": "markdown", - "id": "2cafcd3f", + "id": "6842c1c9", "metadata": { "editable": true }, @@ -3053,7 +3072,7 @@ }, { "cell_type": "markdown", - "id": "b19e7786", + "id": "f7ec748c", "metadata": { "editable": true }, @@ -3071,7 +3090,7 @@ }, { "cell_type": "markdown", - "id": "e046b105", + "id": "6825e389", "metadata": { "editable": true }, @@ -3084,7 +3103,7 @@ }, { "cell_type": "markdown", - "id": "6b9c9c65", + "id": "9328a576", "metadata": { "editable": true }, @@ -3096,7 +3115,7 @@ }, { "cell_type": "markdown", - "id": "9bf87b6d", + "id": "09c73ed7", "metadata": { "editable": true }, @@ -3114,7 +3133,7 @@ }, { "cell_type": "markdown", - "id": "3dec19e7", + "id": "fc6d52ac", "metadata": { "editable": true }, @@ -3132,7 +3151,7 @@ }, { "cell_type": "markdown", - "id": "cfdcc3b8", + "id": "b1a7a657", "metadata": { "editable": true }, @@ -3142,7 +3161,7 @@ }, { "cell_type": "markdown", - "id": "2af74092", + "id": "893ac9c1", "metadata": { "editable": true }, @@ -3160,7 +3179,7 @@ }, { "cell_type": "markdown", - "id": "1267a212", + "id": "f94f743c", "metadata": { "editable": true }, @@ -3179,7 +3198,7 @@ }, { "cell_type": "markdown", - "id": "0f3d9463", + "id": "df1ff4fb", "metadata": { "editable": true }, @@ -3192,7 +3211,7 @@ }, { "cell_type": "markdown", - "id": "093b99c4", + "id": "604facb3", "metadata": { "editable": true }, @@ -3210,7 +3229,7 @@ }, { "cell_type": "markdown", - "id": "89c11305", + "id": "296c4ec9", "metadata": { "editable": true }, @@ -3221,7 +3240,7 @@ }, { "cell_type": "markdown", - "id": "257d8d52", + "id": "3b1937c5", "metadata": { "editable": true }, @@ -3240,7 +3259,7 @@ }, { "cell_type": "markdown", - "id": "c957ad36", + "id": "98b77f08", "metadata": { "editable": true }, @@ -3258,7 +3277,7 @@ }, { "cell_type": "markdown", - "id": "68e73105", + "id": "e72efbc7", "metadata": { "editable": true }, @@ -3271,7 +3290,7 @@ }, { "cell_type": "markdown", - "id": "ac888726", + "id": "d318bb05", "metadata": { "editable": true }, @@ -3291,7 +3310,7 @@ }, { "cell_type": "markdown", - "id": "69e4496b", + "id": "0583bf30", "metadata": { "editable": true }, @@ -3324,7 +3343,7 @@ }, { "cell_type": "markdown", - "id": "2dba373c", + "id": "af1aa099", "metadata": { "editable": true }, @@ -3336,7 +3355,7 @@ }, { "cell_type": "markdown", - "id": "223c3b58", + "id": "bf11e417", "metadata": { "editable": true }, @@ -3355,7 +3374,7 @@ }, { "cell_type": "markdown", - "id": "5900ff0c", + "id": "41f9cca8", "metadata": { "editable": true }, @@ -3369,7 +3388,7 @@ }, { "cell_type": "markdown", - "id": "d1f06df5", + "id": "fa1572f6", "metadata": { "editable": true }, @@ -3392,7 +3411,7 @@ }, { "cell_type": "markdown", - "id": "97f5a7c0", + "id": "a80204ca", "metadata": { "editable": true }, @@ -3411,7 +3430,7 @@ }, { "cell_type": "markdown", - "id": "10efb7e9", + "id": "b2681de5", "metadata": { "editable": true }, @@ -3423,7 +3442,7 @@ }, { "cell_type": "markdown", - "id": "141368fc", + "id": "27d27671", "metadata": { "editable": true }, @@ -3433,7 +3452,7 @@ }, { "cell_type": "markdown", - "id": "d6e68e3f", + "id": "f32c726f", "metadata": { "editable": true }, @@ -3445,7 +3464,7 @@ }, { "cell_type": "markdown", - "id": "f06e702e", + "id": "4090572e", "metadata": { "editable": true }, @@ -3462,7 +3481,7 @@ { "cell_type": "code", "execution_count": 30, - "id": "00e59427", + "id": "6a3ec288", "metadata": { "collapsed": false, "editable": true diff --git a/doc/src/week40/week40.do.txt b/doc/src/week40/week40.do.txt index db3e7bce2..71bc6a87d 100644 --- a/doc/src/week40/week40.do.txt +++ b/doc/src/week40/week40.do.txt @@ -1,6 +1,6 @@ TITLE: Week 40: Gradient descent methods (continued) and start Neural networks AUTHOR: Morten Hjorth-Jensen {copyright, 1999-present|CC BY-NC} at Department of Physics, University of Oslo, Norway & Department of Physics and Astronomy and Facility for Rare Ion Beams, Michigan State University, USA -DATE: October 2-6, 2023 +DATE: September 30-October 4, 2024 @@ -8,32 +8,37 @@ DATE: October 2-6, 2023 ===== Plans for week 40 ===== +!split +===== Lecture Monday September 30, 2024 ===== +!bblock + o Stochastic Gradient descent with examples and automatic differentiation + o If we get time, we start with the basics of Neural Networks, setting up the basic steps, from the simple perceptron model to the multi-layer perceptron model +# * "Video of lecture":"https://youtu.be/75pr3hKY20U" +# * "Whiteboard notes at URL:"https://github.com/CompPhysics/MachineLearning/blob/master/doc/HandWrittenNotes/2023/NotesOct5.pdf" +!eblock + +!split +===== Suggested readings and videos ===== +!bblock Readings and Videos: + o The lecture notes for week 40 (these notes) + o For a good discussion on gradient methods, we would like to recommend Goodfellow et al section 4.3-4.5 and sections 8.3-8.6. We will come back to the latter chapter in our discussion of Neural networks as well. + o For neural networks we recommend Goodfellow et al chapter 6 and Raschka et al pages 48-60 + o Video on gradient descent at URL:"https://www.youtube.com/watch?v=sDv4f4s2SB8" + o Video on stochastic gradient descent at URL:"https://www.youtube.com/watch?v=vMh0zPT0tLI" + o Neural Networks demystified at URL:"https://www.youtube.com/watch?v=bxe2T-V8XRs&list=PLiaHhY2iBX9hdHaRr6b7XevZtgZRa1PoU&ab_channel=WelchLabs" + o Building Neural Networks from scratch at URL:https://www.youtube.com/watch?v=Wo5dMEP_BbI&list=PLQVvvaa0QuDcjD5BAw2DxE6OF2tius3V3&ab_channel=sentdex" +!eblock + +!split +===== Lab sessions Tuesday and Wednesday ===== !bblock Material for the active learning sessions on Tuesday and Wednesday * Work on project 1 and discussions on how to structure your report * No weekly exercises for week 40, project work only - * "Video on how to write scientific reports recorded during one of the lab sessions":"https://youtu.be/tVW1ZDmZnwM" + * Video on how to write scientific reports recorded during one of the lab sessions at URL:"https://youtu.be/tVW1ZDmZnwM" * A general guideline can be found at URL:"https://github.com/CompPhysics/MachineLearning/blob/master/doc/Projects/EvaluationGrading/EvaluationForm.md". !eblock -!bblock Material for the lecture on Thursday October 5, 2023 - * Stochastic Gradient descent with examples and automatic differentiation - * Neural Networks, setting up the basic steps, from the simple perceptron model to the multi-layer perceptron model. - * "Video of lecture":"https://youtu.be/75pr3hKY20U" - * "Whiteboard notes at URL:"https://github.com/CompPhysics/MachineLearning/blob/master/doc/HandWrittenNotes/2023/NotesOct5.pdf" - * Readings and Videos: - * These lecture notes - * For a good discussion on gradient methods, we would like to recommend Goodfellow et al section 4.3-4.5 and sections 8.3-8.6. We will come back to the latter chapter in our discussion of Neural networks as well. - * "Aurelien Geron's chapter 4 on stochastic gradient descent":"https://github.com/CompPhysics/MachineLearning/blob/master/doc/Textbooks/TensorflowML.pdf" - * For neural networks we recommend Goodfellow et al chapter 6. - * "Video on gradient descent":"https://www.youtube.com/watch?v=sDv4f4s2SB8" - * "Video on stochastic gradient descent":"https://www.youtube.com/watch?v=vMh0zPT0tLI" - * "Neural Networks demystified":"https://www.youtube.com/watch?v=bxe2T-V8XRs&list=PLiaHhY2iBX9hdHaRr6b7XevZtgZRa1PoU&ab_channel=WelchLabs" - * "Building Neural Networks from scratch":"https://www.youtube.com/watch?v=Wo5dMEP_BbI&list=PLQVvvaa0QuDcjD5BAw2DxE6OF2tius3V3&ab_channel=sentdex" -!eblock - - - !split ===== Summary from last week, using gradient descent methods, limitations ===== @@ -489,7 +494,7 @@ adaptively change the step size to match the landscape without paying the steep computational price of calculating or approximating Hessians. -Recently, a number of methods have been introduced that accomplish +During the last decade a number of methods have been introduced that accomplish this by tracking not only the gradient, but also the second moment of the gradient. These methods include AdaGrad, AdaDelta, Root Mean Squared Propagation (RMS-Prop), and "ADAM":"https://arxiv.org/abs/1412.6980".