diff --git a/doc/pub/Splines/html/._Splines-bs000.html b/doc/pub/Splines/html/._Splines-bs000.html index 09897dcec..e3f254239 100644 --- a/doc/pub/Splines/html/._Splines-bs000.html +++ b/doc/pub/Splines/html/._Splines-bs000.html @@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -128,18 +128,18 @@ MathJax.Hub.Config({
  • More on Steepest descent
  • The ideal
  • The sensitiveness of the gradient descent
  • -
  • Gradient Descent Example
  • -
  • And a corresponding example using scikit-learn
  • -
  • Convex functions
  • -
  • Convex function
  • -
  • Conditions on convex functions
  • -
  • More on convex functions
  • -
  • Some simple problems
  • -
  • Revisiting our first homework
  • -
  • Gradient descent example
  • -
  • The derivative of the cost/loss function
  • -
  • The Hessian matrix
  • -
  • Simple program
  • +
  • Convex functions
  • +
  • Convex function
  • +
  • Conditions on convex functions
  • +
  • More on convex functions
  • +
  • Some simple problems
  • +
  • Revisiting our first homework
  • +
  • Gradient descent example
  • +
  • The derivative of the cost/loss function
  • +
  • The Hessian matrix
  • +
  • Simple program
  • +
  • Gradient Descent Example
  • +
  • And a corresponding example using scikit-learn
  • Gradient descent and Ridge
  • Stochastic Gradient Descent
  • Computation of gradients
  • @@ -193,7 +193,7 @@ MathJax.Hub.Config({
    [2] Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University

    -

    Sep 20, 2018

    +

    Sep 21, 2018


    diff --git a/doc/pub/Splines/html/._Splines-bs001.html b/doc/pub/Splines/html/._Splines-bs001.html index ad05fabe2..fa2e9e8ee 100644 --- a/doc/pub/Splines/html/._Splines-bs001.html +++ b/doc/pub/Splines/html/._Splines-bs001.html @@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -128,18 +128,18 @@ MathJax.Hub.Config({

  • More on Steepest descent
  • The ideal
  • The sensitiveness of the gradient descent
  • -
  • Gradient Descent Example
  • -
  • And a corresponding example using scikit-learn
  • -
  • Convex functions
  • -
  • Convex function
  • -
  • Conditions on convex functions
  • -
  • More on convex functions
  • -
  • Some simple problems
  • -
  • Revisiting our first homework
  • -
  • Gradient descent example
  • -
  • The derivative of the cost/loss function
  • -
  • The Hessian matrix
  • -
  • Simple program
  • +
  • Convex functions
  • +
  • Convex function
  • +
  • Conditions on convex functions
  • +
  • More on convex functions
  • +
  • Some simple problems
  • +
  • Revisiting our first homework
  • +
  • Gradient descent example
  • +
  • The derivative of the cost/loss function
  • +
  • The Hessian matrix
  • +
  • Simple program
  • +
  • Gradient Descent Example
  • +
  • And a corresponding example using scikit-learn
  • Gradient descent and Ridge
  • Stochastic Gradient Descent
  • Computation of gradients
  • diff --git a/doc/pub/Splines/html/._Splines-bs002.html b/doc/pub/Splines/html/._Splines-bs002.html index 2c97de8df..034502d3a 100644 --- a/doc/pub/Splines/html/._Splines-bs002.html +++ b/doc/pub/Splines/html/._Splines-bs002.html @@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -128,18 +128,18 @@ MathJax.Hub.Config({
  • More on Steepest descent
  • The ideal
  • The sensitiveness of the gradient descent
  • -
  • Gradient Descent Example
  • -
  • And a corresponding example using scikit-learn
  • -
  • Convex functions
  • -
  • Convex function
  • -
  • Conditions on convex functions
  • -
  • More on convex functions
  • -
  • Some simple problems
  • -
  • Revisiting our first homework
  • -
  • Gradient descent example
  • -
  • The derivative of the cost/loss function
  • -
  • The Hessian matrix
  • -
  • Simple program
  • +
  • Convex functions
  • +
  • Convex function
  • +
  • Conditions on convex functions
  • +
  • More on convex functions
  • +
  • Some simple problems
  • +
  • Revisiting our first homework
  • +
  • Gradient descent example
  • +
  • The derivative of the cost/loss function
  • +
  • The Hessian matrix
  • +
  • Simple program
  • +
  • Gradient Descent Example
  • +
  • And a corresponding example using scikit-learn
  • Gradient descent and Ridge
  • Stochastic Gradient Descent
  • Computation of gradients
  • @@ -185,11 +185,13 @@ direction of the negative gradient \( -\nabla F(\mathbf{x}) \).

    It can be shown that if $$ -\mathbf{x}_{k+1} = \mathbf{x}_k - \gamma_k \nabla F(\mathbf{x}_k), \ \ \gamma_k > 0 +\mathbf{x}_{k+1} = \mathbf{x}_k - \gamma_k \nabla F(\mathbf{x}_k), $$ +with \( \gamma_k > 0 \). +

    -for \( \gamma_k \) small enough, then \( F(\mathbf{x}_{k+1}) \leq +For \( \gamma_k \) small enough, then \( F(\mathbf{x}_{k+1}) \leq F(\mathbf{x}_k) \). This means that for a sufficiently small \( \gamma_k \) we are always moving towards smaller function values, i.e a minimum. diff --git a/doc/pub/Splines/html/._Splines-bs003.html b/doc/pub/Splines/html/._Splines-bs003.html index 90b01b5c8..1537aadc8 100644 --- a/doc/pub/Splines/html/._Splines-bs003.html +++ b/doc/pub/Splines/html/._Splines-bs003.html @@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -128,18 +128,18 @@ MathJax.Hub.Config({

  • More on Steepest descent
  • The ideal
  • The sensitiveness of the gradient descent
  • -
  • Gradient Descent Example
  • -
  • And a corresponding example using scikit-learn
  • -
  • Convex functions
  • -
  • Convex function
  • -
  • Conditions on convex functions
  • -
  • More on convex functions
  • -
  • Some simple problems
  • -
  • Revisiting our first homework
  • -
  • Gradient descent example
  • -
  • The derivative of the cost/loss function
  • -
  • The Hessian matrix
  • -
  • Simple program
  • +
  • Convex functions
  • +
  • Convex function
  • +
  • Conditions on convex functions
  • +
  • More on convex functions
  • +
  • Some simple problems
  • +
  • Revisiting our first homework
  • +
  • Gradient descent example
  • +
  • The derivative of the cost/loss function
  • +
  • The Hessian matrix
  • +
  • Simple program
  • +
  • Gradient Descent Example
  • +
  • And a corresponding example using scikit-learn
  • Gradient descent and Ridge
  • Stochastic Gradient Descent
  • Computation of gradients
  • diff --git a/doc/pub/Splines/html/._Splines-bs004.html b/doc/pub/Splines/html/._Splines-bs004.html index 9484ecc0c..47ec80c33 100644 --- a/doc/pub/Splines/html/._Splines-bs004.html +++ b/doc/pub/Splines/html/._Splines-bs004.html @@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -128,18 +128,18 @@ MathJax.Hub.Config({
  • More on Steepest descent
  • The ideal
  • The sensitiveness of the gradient descent
  • -
  • Gradient Descent Example
  • -
  • And a corresponding example using scikit-learn
  • -
  • Convex functions
  • -
  • Convex function
  • -
  • Conditions on convex functions
  • -
  • More on convex functions
  • -
  • Some simple problems
  • -
  • Revisiting our first homework
  • -
  • Gradient descent example
  • -
  • The derivative of the cost/loss function
  • -
  • The Hessian matrix
  • -
  • Simple program
  • +
  • Convex functions
  • +
  • Convex function
  • +
  • Conditions on convex functions
  • +
  • More on convex functions
  • +
  • Some simple problems
  • +
  • Revisiting our first homework
  • +
  • Gradient descent example
  • +
  • The derivative of the cost/loss function
  • +
  • The Hessian matrix
  • +
  • Simple program
  • +
  • Gradient Descent Example
  • +
  • And a corresponding example using scikit-learn
  • Gradient descent and Ridge
  • Stochastic Gradient Descent
  • Computation of gradients
  • @@ -177,7 +177,7 @@ MathJax.Hub.Config({

    The ideal

    -Ideally the sequence \( \{ \mathbf{x}_k \}_{k=0} \) converges to a global +Ideally the sequence \( \{\mathbf{x}_k \}_{k=0} \) converges to a global minimum of the function \( F \). In general we do not know if we are in a global or local minimum. In the special case when \( F \) is a convex function, all local minima are also global minima, so in this case diff --git a/doc/pub/Splines/html/._Splines-bs005.html b/doc/pub/Splines/html/._Splines-bs005.html index f9a3208bd..a4e58bab5 100644 --- a/doc/pub/Splines/html/._Splines-bs005.html +++ b/doc/pub/Splines/html/._Splines-bs005.html @@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -128,18 +128,18 @@ MathJax.Hub.Config({

  • More on Steepest descent
  • The ideal
  • The sensitiveness of the gradient descent
  • -
  • Gradient Descent Example
  • -
  • And a corresponding example using scikit-learn
  • -
  • Convex functions
  • -
  • Convex function
  • -
  • Conditions on convex functions
  • -
  • More on convex functions
  • -
  • Some simple problems
  • -
  • Revisiting our first homework
  • -
  • Gradient descent example
  • -
  • The derivative of the cost/loss function
  • -
  • The Hessian matrix
  • -
  • Simple program
  • +
  • Convex functions
  • +
  • Convex function
  • +
  • Conditions on convex functions
  • +
  • More on convex functions
  • +
  • Some simple problems
  • +
  • Revisiting our first homework
  • +
  • Gradient descent example
  • +
  • The derivative of the cost/loss function
  • +
  • The Hessian matrix
  • +
  • Simple program
  • +
  • Gradient Descent Example
  • +
  • And a corresponding example using scikit-learn
  • Gradient descent and Ridge
  • Stochastic Gradient Descent
  • Computation of gradients
  • @@ -177,7 +177,8 @@ MathJax.Hub.Config({

    The sensitiveness of the gradient descent

    -GD is sensitive to the choice of learning rate \( \gamma_k \). This is due +The gradient descent method +is sensitive to the choice of learning rate \( \gamma_k \). This is due to the fact that we are only guaranteed that \( F(\mathbf{x}_{k+1}) \leq F(\mathbf{x}_k) \) for sufficiently small \( \gamma_k \). The problem is to determine an optimal learning rate. If the learning rate is chosen too diff --git a/doc/pub/Splines/html/._Splines-bs006.html b/doc/pub/Splines/html/._Splines-bs006.html index 306b99cdb..1ba9ffe66 100644 --- a/doc/pub/Splines/html/._Splines-bs006.html +++ b/doc/pub/Splines/html/._Splines-bs006.html @@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -128,18 +128,18 @@ MathJax.Hub.Config({

  • More on Steepest descent
  • The ideal
  • The sensitiveness of the gradient descent
  • -
  • Gradient Descent Example
  • -
  • And a corresponding example using scikit-learn
  • -
  • Convex functions
  • -
  • Convex function
  • -
  • Conditions on convex functions
  • -
  • More on convex functions
  • -
  • Some simple problems
  • -
  • Revisiting our first homework
  • -
  • Gradient descent example
  • -
  • The derivative of the cost/loss function
  • -
  • The Hessian matrix
  • -
  • Simple program
  • +
  • Convex functions
  • +
  • Convex function
  • +
  • Conditions on convex functions
  • +
  • More on convex functions
  • +
  • Some simple problems
  • +
  • Revisiting our first homework
  • +
  • Gradient descent example
  • +
  • The derivative of the cost/loss function
  • +
  • The Hessian matrix
  • +
  • Simple program
  • +
  • Gradient Descent Example
  • +
  • And a corresponding example using scikit-learn
  • Gradient descent and Ridge
  • Stochastic Gradient Descent
  • Computation of gradients
  • @@ -172,54 +172,25 @@ MathJax.Hub.Config({

     

     

     

    - + -

    Gradient Descent Example

    +

    Convex functions

    -We revisit now our simple linear regression example with a linear polynomial. +Ideally we want our cost/loss function to be convex(concave). +

    +First we give the definition of a convex set: A set \( C \) in +\( \mathbb{R}^n \) is said to be convex if, for all \( x \) and \( y \) in \( C \) and +all \( t \in (0,1) \) , the point \( (1 − t)x + ty \) also belongs to +C. Geometrically this means that every point on the line segment +connecting \( x \) and \( y \) is in \( C \) as discussed below. - -

    # Importing various packages
    -from random import random, seed
    -import numpy as np
    -import matplotlib.pyplot as plt
    -from mpl_toolkits.mplot3d import Axes3D
    -from matplotlib import cm
    -from matplotlib.ticker import LinearLocator, FormatStrFormatter
    -import sys
    +

    +The convex subsets of \( \mathbb{R} \) are the intervals of +\( \mathbb{R} \). Examples of convex sets of \( \mathbb{R}^2 \) are the +regular polygons (triangles, rectangles, pentagons, etc...). -x = 2*np.random.rand(100,1) -y = 4+3*x+np.random.randn(100,1) - -xb = np.c_[np.ones((100,1)), x] -beta_linreg = np.linalg.inv(xb.T.dot(xb)).dot(xb.T).dot(y) -print(beta_linreg) -beta = np.random.randn(2,1) - -eta = 0.1 -Niterations = 1000 -m = 100 - -for iter in range(Niterations): - gradients = 2.0/m*xb.T.dot(xb.dot(beta)-y) - beta -= eta*gradients - -print(beta) -xnew = np.array([[0],[2]]) -xbnew = np.c_[np.ones((2,1)), xnew] -ypredict = xbnew.dot(beta) -ypredict2 = xbnew.dot(beta_linreg) -plt.plot(xnew, ypredict, "r-") -plt.plot(xnew, ypredict2, "b-") -plt.plot(x, y ,'ro') -plt.axis([0,2.0,0, 15.0]) -plt.xlabel(r'$x$') -plt.ylabel(r'$y$') -plt.title(r'Gradient descent example') -plt.show() -

    diff --git a/doc/pub/Splines/html/._Splines-bs007.html b/doc/pub/Splines/html/._Splines-bs007.html index 6b2852c4c..57c477e77 100644 --- a/doc/pub/Splines/html/._Splines-bs007.html +++ b/doc/pub/Splines/html/._Splines-bs007.html @@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -128,18 +128,18 @@ MathJax.Hub.Config({

  • More on Steepest descent
  • The ideal
  • The sensitiveness of the gradient descent
  • -
  • Gradient Descent Example
  • -
  • And a corresponding example using scikit-learn
  • -
  • Convex functions
  • -
  • Convex function
  • -
  • Conditions on convex functions
  • -
  • More on convex functions
  • -
  • Some simple problems
  • -
  • Revisiting our first homework
  • -
  • Gradient descent example
  • -
  • The derivative of the cost/loss function
  • -
  • The Hessian matrix
  • -
  • Simple program
  • +
  • Convex functions
  • +
  • Convex function
  • +
  • Conditions on convex functions
  • +
  • More on convex functions
  • +
  • Some simple problems
  • +
  • Revisiting our first homework
  • +
  • Gradient descent example
  • +
  • The derivative of the cost/loss function
  • +
  • The Hessian matrix
  • +
  • Simple program
  • +
  • Gradient Descent Example
  • +
  • And a corresponding example using scikit-learn
  • Gradient descent and Ridge
  • Stochastic Gradient Descent
  • Computation of gradients
  • @@ -174,27 +174,11 @@ MathJax.Hub.Config({ -

    And a corresponding example using scikit-learn

    +

    Convex function

    +Convex function: Let \( X \subset \mathbb{R}^n \) be a convex set. Assume that the function \( f: X \rightarrow \mathbb{R} \) is continuous, then \( f \) is said to be convex if $$f(tx_1 + (1-t)x_2) \leq tf(x_1) + (1-t)f(x_2) $$ for all \( x_1, x_2 \in X \) and for all \( t \in [0,1] \). If \( \leq \) is replaced with a strict inequaltiy in the definition, we demand \( x_1 \neq x_2 \) and \( t\in(0,1) \) then \( f \) is said to be strictly convex. For a single variable function, convexity means that if you draw a straight line connecting \( f(x_1) \) and \( f(x_2) \), the value of the function on the interval \( [x_1,x_2] \) is always below the line as illustrated below. - -

    # Importing various packages
    -from random import random, seed
    -import numpy as np
    -import matplotlib.pyplot as plt
    -from sklearn.linear_model import SGDRegressor
    -
    -x = 2*np.random.rand(100,1)
    -y = 4+3*x+np.random.randn(100,1)
    -
    -xb = np.c_[np.ones((100,1)), x]
    -beta_linreg = np.linalg.inv(xb.T.dot(xb)).dot(xb.T).dot(y)
    -print(beta_linreg)
    -sgdreg = SGDRegressor(n_iter = 50, penalty=None, eta0=0.1)
    -sgdreg.fit(x,y.ravel())
    -print(sgdreg.intercept_, sgdreg.coef_)
    -

    diff --git a/doc/pub/Splines/html/._Splines-bs008.html b/doc/pub/Splines/html/._Splines-bs008.html index 86e9ac845..f0565b9ad 100644 --- a/doc/pub/Splines/html/._Splines-bs008.html +++ b/doc/pub/Splines/html/._Splines-bs008.html @@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -128,18 +128,18 @@ MathJax.Hub.Config({

  • More on Steepest descent
  • The ideal
  • The sensitiveness of the gradient descent
  • -
  • Gradient Descent Example
  • -
  • And a corresponding example using scikit-learn
  • -
  • Convex functions
  • -
  • Convex function
  • -
  • Conditions on convex functions
  • -
  • More on convex functions
  • -
  • Some simple problems
  • -
  • Revisiting our first homework
  • -
  • Gradient descent example
  • -
  • The derivative of the cost/loss function
  • -
  • The Hessian matrix
  • -
  • Simple program
  • +
  • Convex functions
  • +
  • Convex function
  • +
  • Conditions on convex functions
  • +
  • More on convex functions
  • +
  • Some simple problems
  • +
  • Revisiting our first homework
  • +
  • Gradient descent example
  • +
  • The derivative of the cost/loss function
  • +
  • The Hessian matrix
  • +
  • Simple program
  • +
  • Gradient Descent Example
  • +
  • And a corresponding example using scikit-learn
  • Gradient descent and Ridge
  • Stochastic Gradient Descent
  • Computation of gradients
  • @@ -172,24 +172,48 @@ MathJax.Hub.Config({

     

     

     

    - + -

    Convex functions

    +

    Conditions on convex functions

    -Ideally we want our cost/loss function to be convex(concave). +In the following we state first and second-order conditions which +ensures convexity of a function \( f \). We write \( D_f \) to denote the +domain of \( f \), i.e the subset of \( R^n \) where \( f \) is defined. For more +details and proofs we refer to: S. Boyd and L. Vandenberghe. Convex Optimization. Cambridge University Press.

    -First we give the definition of a convex set: A set \( C \) in -\( \mathbb{R}^n \) is said to be convex if, for all \( x \) and \( y \) in \( C \) and -all \( t \in (0,1) \) , the point \( (1 − t)x + ty \) also belongs to -C. Geometrically this means that every point on the line segment -connecting \( x \) and \( y \) is in \( C \) as discussed below. +

    +
    +

    +Suppose \( f \) is differentiable (i.e \( \nabla f(x) \) is well defined for +all \( x \) in the domain of \( f \)). Then \( f \) is convex if and only if \( D_f \) +is a convex set and $$f(y) \geq f(x) + \nabla f(x)^T (y-x) $$ holds +for all \( x,y \in D_f \). This condition means that for a convex function +the first order Taylor expansion (right hand side above) at any point +a global under estimator of the function. To convince yourself you can +make a drawing of \( f(x) = x^2+1 \) and draw the tangent line to \( f(x) \) and +note that it is always below the graph. +

    +
    +

    -The convex subsets of \( \mathbb{R} \) are the intervals of -\( \mathbb{R} \). Examples of convex sets of \( \mathbb{R}^2 \) are the -regular polygons (triangles, rectangles, pentagons, etc...). +

    +
    +

    +Assume that \( f \) is twice +differentiable, i.e the Hessian matrix exists at each point in +\( D_f \). Then \( f \) is convex if and only if \( D_f \) is a convex set and its +Hessian is positive semi-definite for all \( x\in D_f \). For a +single-variable function this reduces to \( f''(x) \geq 0 \). Geometrically this means that \( f \) has nonnegative curvature +everywhere. +

    +
    + + +

    +This condition is particularly useful since it gives us an procedure for determining if the function under consideration is convex, apart from using the definition.

    diff --git a/doc/pub/Splines/html/._Splines-bs009.html b/doc/pub/Splines/html/._Splines-bs009.html index 7ae79d754..15e45e732 100644 --- a/doc/pub/Splines/html/._Splines-bs009.html +++ b/doc/pub/Splines/html/._Splines-bs009.html @@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -128,18 +128,18 @@ MathJax.Hub.Config({

  • More on Steepest descent
  • The ideal
  • The sensitiveness of the gradient descent
  • -
  • Gradient Descent Example
  • -
  • And a corresponding example using scikit-learn
  • -
  • Convex functions
  • -
  • Convex function
  • -
  • Conditions on convex functions
  • -
  • More on convex functions
  • -
  • Some simple problems
  • -
  • Revisiting our first homework
  • -
  • Gradient descent example
  • -
  • The derivative of the cost/loss function
  • -
  • The Hessian matrix
  • -
  • Simple program
  • +
  • Convex functions
  • +
  • Convex function
  • +
  • Conditions on convex functions
  • +
  • More on convex functions
  • +
  • Some simple problems
  • +
  • Revisiting our first homework
  • +
  • Gradient descent example
  • +
  • The derivative of the cost/loss function
  • +
  • The Hessian matrix
  • +
  • Simple program
  • +
  • Gradient Descent Example
  • +
  • And a corresponding example using scikit-learn
  • Gradient descent and Ridge
  • Stochastic Gradient Descent
  • Computation of gradients
  • @@ -174,10 +174,33 @@ MathJax.Hub.Config({ -

    Convex function

    +

    More on convex functions

    -Convex function: Let \( X \subset \mathbb{R}^n \) be a convex set. Assume that the function \( f: X \rightarrow \mathbb{R} \) is continuous, then \( f \) is said to be convex if $$f(tx_1 + (1-t)x_2) \leq tf(x_1) + (1-t)f(x_2) $$ for all \( x_1, x_2 \in X \) and for all \( t \in [0,1] \). If \( \leq \) is replaced with a strict inequaltiy in the definition, we demand \( x_1 \neq x_2 \) and \( t\in(0,1) \) then \( f \) is said to be strictly convex. For a single variable function, convexity means that if you draw a straight line connecting \( f(x_1) \) and \( f(x_2) \), the value of the function on the interval \( [x_1,x_2] \) is always below the line as illustrated below. +The next result is of great importance to us and the reason why we are +going on about convex functions. In machine learning we frequently +have to minimize a loss/cost function in order to find the best +parameters for the model we are considering. + +

    +Ideally we want the +global minimum (for high-dimensional models it is hard to know +if we have local or global minimum). However, if the cost/loss function +is convex the following result provides invaluable information: + +

    +

    +
    +

    +Consider the problem of finding \( x \in \mathbb{R}^n \) such that \( f(x) \) +is minimal, where \( f \) is convex and differentiable. Then, any point +\( x^* \) that satisfies \( \nabla f(x^*) = 0 \) is a global minimum. +

    +
    + + +

    +This result means that if we know that the cost/loss function is convex and we are able to find a minimum, we are guaranteed that it is a global minimum.

    diff --git a/doc/pub/Splines/html/._Splines-bs010.html b/doc/pub/Splines/html/._Splines-bs010.html index c1bb98aa7..91bd8be2f 100644 --- a/doc/pub/Splines/html/._Splines-bs010.html +++ b/doc/pub/Splines/html/._Splines-bs010.html @@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -128,18 +128,18 @@ MathJax.Hub.Config({

  • More on Steepest descent
  • The ideal
  • The sensitiveness of the gradient descent
  • -
  • Gradient Descent Example
  • -
  • And a corresponding example using scikit-learn
  • -
  • Convex functions
  • -
  • Convex function
  • -
  • Conditions on convex functions
  • -
  • More on convex functions
  • -
  • Some simple problems
  • -
  • Revisiting our first homework
  • -
  • Gradient descent example
  • -
  • The derivative of the cost/loss function
  • -
  • The Hessian matrix
  • -
  • Simple program
  • +
  • Convex functions
  • +
  • Convex function
  • +
  • Conditions on convex functions
  • +
  • More on convex functions
  • +
  • Some simple problems
  • +
  • Revisiting our first homework
  • +
  • Gradient descent example
  • +
  • The derivative of the cost/loss function
  • +
  • The Hessian matrix
  • +
  • Simple program
  • +
  • Gradient Descent Example
  • +
  • And a corresponding example using scikit-learn
  • Gradient descent and Ridge
  • Stochastic Gradient Descent
  • Computation of gradients
  • @@ -174,47 +174,29 @@ MathJax.Hub.Config({ -

    Conditions on convex functions

    +

    Some simple problems

    -

    -In the following we state first and second-order conditions which -ensures convexity of a function \( f \). We write \( D_f \) to denote the -domain of \( f \), i.e the subset of \( R^n \) where \( f \) is defined. For more -details and proofs we refer to: S. Boyd and L. Vandenberghe. Convex Optimization. Cambridge University Press. +

      +
    1. Show that \( f(x)=x^2 \) is convex for \( x \in \mathbb{R} \) using the definition of convexity. Hint: If you re-write the definition, \( f \) is convex if the following holds for all \( x,y \in D_f \) and any \( \lambda \in [0,1] \) $\lambda f(x)+(1-\lambda)f(y)-f(\lambda x + (1-\lambda) y ) \geq 0$.
    2. +
    3. Using the second order condition show that the following functions are convex on the specified domain.
    4. -

      -

      -
      -

      -Suppose \( f \) is differentiable (i.e \( \nabla f(x) \) is well defined for -all \( x \) in the domain of \( f \)). Then \( f \) is convex if and only if \( D_f \) -is a convex set and $$f(y) \geq f(x) + \nabla f(x)^T (y-x) $$ holds -for all \( x,y \in D_f \). This condition means that for a convex function -the first order Taylor expansion (right hand side above) at any point -a global under estimator of the function. To convince yourself you can -make a drawing of f(x) = x^2+1 and draw the tangent line to \( f(x) \) and -note that it is always below the graph. -

      -
      + +
    5. Let \( f(x) = x^2 \) and \( g(x) = e^x \). Show that \( f(g(x)) \) and \( g(f(x)) \) is convex for \( x \in \mathbb{R} \). Also show that if \( f(x) \) is any convex function than \( h(x) = e^{f(x)} \) is convex.
    6. +
    7. A norm is any function that satisfy the following properties
    8. -

      -

      -
      -

      -Assume that \( f \) is twice -differentiable, i.e the Hessian matrix exists at each point in -\( D_f \). Then \( f \) is convex if and only if \( D_f \) is a convex set and its -Hessian is positive semi-definite for all \( x\in D_f \). For a -single-variable function this reduces to \( f''(x) \geq -0 \). Geometrically this means that \( f \) has nonnegative curvature -everywhere. -

      -
      + +
    -

    -This condition is particularly useful since it gives us an procedure for determining if the function under consideration is convex, apart from using the definition. +Using the definition of convexity, try to show that a function satisfying the properties above is convex (the third condition is not needed to show this).

    diff --git a/doc/pub/Splines/html/._Splines-bs011.html b/doc/pub/Splines/html/._Splines-bs011.html index 8a0925577..fc3bba3df 100644 --- a/doc/pub/Splines/html/._Splines-bs011.html +++ b/doc/pub/Splines/html/._Splines-bs011.html @@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -128,18 +128,18 @@ MathJax.Hub.Config({

  • More on Steepest descent
  • The ideal
  • The sensitiveness of the gradient descent
  • -
  • Gradient Descent Example
  • -
  • And a corresponding example using scikit-learn
  • -
  • Convex functions
  • -
  • Convex function
  • -
  • Conditions on convex functions
  • -
  • More on convex functions
  • -
  • Some simple problems
  • -
  • Revisiting our first homework
  • -
  • Gradient descent example
  • -
  • The derivative of the cost/loss function
  • -
  • The Hessian matrix
  • -
  • Simple program
  • +
  • Convex functions
  • +
  • Convex function
  • +
  • Conditions on convex functions
  • +
  • More on convex functions
  • +
  • Some simple problems
  • +
  • Revisiting our first homework
  • +
  • Gradient descent example
  • +
  • The derivative of the cost/loss function
  • +
  • The Hessian matrix
  • +
  • Simple program
  • +
  • Gradient Descent Example
  • +
  • And a corresponding example using scikit-learn
  • Gradient descent and Ridge
  • Stochastic Gradient Descent
  • Computation of gradients
  • @@ -172,35 +172,37 @@ MathJax.Hub.Config({

     

     

     

    - + -

    More on convex functions

    +

    Revisiting our first homework

    -The next result is of great importance to us and the reason why we are -going on about convex functions. In machine learning we frequently -have to minimize a loss/cost function in order to find the best -parameters for the model we are considering. +We will use linear regression as a case study for the gradient descent +methods. Linear regression is a great test case for the gradient +descent methods discussed in the lectures since it has several +desirable properties such as: -

    -Ideally we want the -global minimum (for high-dimensional models it is hard to know -if we have local or global minimum). However, if the cost/loss function -is convex the following result provides invaluable information: +

      +
    1. An analytical solution (recall homework set 1).
    2. +
    3. The gradient can be computed analytically.
    4. +
    5. The cost function is convex which guarantees that gradient descent converges for small enough learning rates
    6. +
    -

    -

    -
    -

    -Consider the problem of finding \( x \in \mathbb{R}^n \) such that \( f(x) \) -is minimal, where \( f \) is convex and differentiable. Then, any point -\( x^* \) that satisfies \( \nabla f(x^*) = 0 \) is a global minimum. -

    -
    +We revisit the example from homework set 1 where we had +$$ +y_i = 5x_i^2 + 0.1\xi_i, \ i=1,\cdots,100 +$$ +with \( x_i \in [0,1] \) chosen randomly with a uniform distribution. Additionally \( \xi_i \) represents stochastic noise chosen according to a normal distribution \( \cal {N}(0,1) \). +The linear regression model is given by +$$ +h_\beta(x) = \hat{y} = \beta_0 + \beta_1 x, +$$ -

    -This result means that if we know that the cost/loss function is convex and we are able to find a minimum, we are guaranteed that it is a global minimum. +such that +$$ +\hat{y}_i = \beta_0 + \beta_1 x_i. +$$

    diff --git a/doc/pub/Splines/html/._Splines-bs012.html b/doc/pub/Splines/html/._Splines-bs012.html index 29a039b11..27137f59b 100644 --- a/doc/pub/Splines/html/._Splines-bs012.html +++ b/doc/pub/Splines/html/._Splines-bs012.html @@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -128,18 +128,18 @@ MathJax.Hub.Config({

  • More on Steepest descent
  • The ideal
  • The sensitiveness of the gradient descent
  • -
  • Gradient Descent Example
  • -
  • And a corresponding example using scikit-learn
  • -
  • Convex functions
  • -
  • Convex function
  • -
  • Conditions on convex functions
  • -
  • More on convex functions
  • -
  • Some simple problems
  • -
  • Revisiting our first homework
  • -
  • Gradient descent example
  • -
  • The derivative of the cost/loss function
  • -
  • The Hessian matrix
  • -
  • Simple program
  • +
  • Convex functions
  • +
  • Convex function
  • +
  • Conditions on convex functions
  • +
  • More on convex functions
  • +
  • Some simple problems
  • +
  • Revisiting our first homework
  • +
  • Gradient descent example
  • +
  • The derivative of the cost/loss function
  • +
  • The Hessian matrix
  • +
  • Simple program
  • +
  • Gradient Descent Example
  • +
  • And a corresponding example using scikit-learn
  • Gradient descent and Ridge
  • Stochastic Gradient Descent
  • Computation of gradients
  • @@ -172,31 +172,29 @@ MathJax.Hub.Config({

     

     

     

    - + -

    Some simple problems

    +

    Gradient descent example

    -
      -
    1. Show that \( f(x)=x^2 \) is convex for \( x \in \mathbb{R} \) using the definition of convexity. Hint: If you re-write the definition, \( f \) is convex if the following holds for all \( x,y \in D_f \) and any \( \lambda \in [0,1] \) $\lambda f(x) + (1-\lambda)f(y) - f(\lambda x + (1-\lambda) y ) \geq 0. $
    2. -
    3. Using the second order condition show that the following functions are convex on the specified domain.
    4. +

      +Let \( \mathbf{y} = (y_1,\cdots,y_n)^T \), \( \mathbf{\hat{y}} = (\hat{y}_1,\cdots,\hat{y}_n)^T \) and \( \beta = (\beta_0, \beta_1)^T \) -

      +

      +It is convenient to write \( \mathbf{\hat{y}} = X\beta \) where \( X \in \mathbb{R}^{100 \times 2} \) is the design matrix given by +$$ +X \equiv \begin{bmatrix} +1 & x_1 \\ +\vdots & \vdots \\ +1 & x_{100} & \\ +\end{bmatrix}. +$$ -

    5. Let \( f(x) = x^2 \) and \( g(x) = e^x \). Show that \( f(g(x)) \) and \( g(f(x)) \) is convex for \( x \in \mathbb{R} \). Also show that if \( f(x) \) is any convex function than \( h(x) = e^{f(x)} \) is convex.
    6. -
    7. A norm is any function that satisfy the following properties
    8. +The loss function is given by +$$ +C(\beta) = ||X\beta-\mathbf{y}||^2 = ||X\beta||^2 - 2 \mathbf{y}^T X\beta + ||\mathbf{y}||^2 = \sum_{i=1}^{100} (\beta_0 + \beta_1 x_i)^2 - 2 y_i (\beta_0 + \beta_1 x_i) + y_i^2 +$$ - - -
    - -Using the definition of convexity, try to show that a function satisfying the properties above is convex (the third condition is not needed to show this). +and we want to find \( \beta \) such that \( C(\beta) \) is minimized.

    diff --git a/doc/pub/Splines/html/._Splines-bs013.html b/doc/pub/Splines/html/._Splines-bs013.html index 2a7103a97..372867cf3 100644 --- a/doc/pub/Splines/html/._Splines-bs013.html +++ b/doc/pub/Splines/html/._Splines-bs013.html @@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -128,18 +128,18 @@ MathJax.Hub.Config({

  • More on Steepest descent
  • The ideal
  • The sensitiveness of the gradient descent
  • -
  • Gradient Descent Example
  • -
  • And a corresponding example using scikit-learn
  • -
  • Convex functions
  • -
  • Convex function
  • -
  • Conditions on convex functions
  • -
  • More on convex functions
  • -
  • Some simple problems
  • -
  • Revisiting our first homework
  • -
  • Gradient descent example
  • -
  • The derivative of the cost/loss function
  • -
  • The Hessian matrix
  • -
  • Simple program
  • +
  • Convex functions
  • +
  • Convex function
  • +
  • Conditions on convex functions
  • +
  • More on convex functions
  • +
  • Some simple problems
  • +
  • Revisiting our first homework
  • +
  • Gradient descent example
  • +
  • The derivative of the cost/loss function
  • +
  • The Hessian matrix
  • +
  • Simple program
  • +
  • Gradient Descent Example
  • +
  • And a corresponding example using scikit-learn
  • Gradient descent and Ridge
  • Stochastic Gradient Descent
  • Computation of gradients
  • @@ -172,37 +172,19 @@ MathJax.Hub.Config({

     

     

     

    - + -

    Revisiting our first homework

    +

    The derivative of the cost/loss function

    -We will use linear regression as a case study for the gradient descent -methods. Linear regression is a great test case for the gradient -descent methods discussed in the lectures since it has several -desirable properties such as: - -

      -
    1. An analytical solution (recall homework set 1).
    2. -
    3. The gradient can be computed analytically.
    4. -
    5. The cost function is convex which guarantees that gradient descent converges for small enough learning rates
    6. -
    - -We revisit the example from homework set 1 where we had +Computing \( \partial C(\beta) / \partial \beta_0 \) and \( \partial C(\beta) / \partial \beta_1 \) we can show that the gradient can be written as $$ -y_i = 5x_i^2 + 0.1\xi_i, \ i=1,\cdots,100 +\nabla_{\beta} C(\beta) = (\partial C(\beta) / \partial \beta_0, \partial C(\beta) / \partial \beta_1)^T = 2\begin{bmatrix} \sum_{i=1}^{100} \left(\beta_0+\beta_1x_i-y_i\right) \\ +\sum_{i=1}^{100}\left( x_i (\beta_0+\beta_1x_i)-y_ix_i\right) \\ +\end{bmatrix} = 2X^T(X\beta - \mathbf{y}), $$ -with \( x_i \in [0,1] \) chosen randomly with a uniform distribution. Additionally \( \xi_i \) represents stochastic noise chosen according to a normal distribution \( \cal {N}(0,1) \). -The linear regression model is given by -$$ -h_\beta(x) = \hat{y} = \beta_0 + \beta_1 x, -$$ - -such that -$$ -\hat{y}_i = \beta_0 + \beta_1 x_i. -$$ +where \( X \) is the design matrix defined above.

    diff --git a/doc/pub/Splines/html/._Splines-bs014.html b/doc/pub/Splines/html/._Splines-bs014.html index 596f81cd4..9374c1ab0 100644 --- a/doc/pub/Splines/html/._Splines-bs014.html +++ b/doc/pub/Splines/html/._Splines-bs014.html @@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -128,18 +128,18 @@ MathJax.Hub.Config({

  • More on Steepest descent
  • The ideal
  • The sensitiveness of the gradient descent
  • -
  • Gradient Descent Example
  • -
  • And a corresponding example using scikit-learn
  • -
  • Convex functions
  • -
  • Convex function
  • -
  • Conditions on convex functions
  • -
  • More on convex functions
  • -
  • Some simple problems
  • -
  • Revisiting our first homework
  • -
  • Gradient descent example
  • -
  • The derivative of the cost/loss function
  • -
  • The Hessian matrix
  • -
  • Simple program
  • +
  • Convex functions
  • +
  • Convex function
  • +
  • Conditions on convex functions
  • +
  • More on convex functions
  • +
  • Some simple problems
  • +
  • Revisiting our first homework
  • +
  • Gradient descent example
  • +
  • The derivative of the cost/loss function
  • +
  • The Hessian matrix
  • +
  • Simple program
  • +
  • Gradient Descent Example
  • +
  • And a corresponding example using scikit-learn
  • Gradient descent and Ridge
  • Stochastic Gradient Descent
  • Computation of gradients
  • @@ -172,32 +172,18 @@ MathJax.Hub.Config({

     

     

     

    - + -

    Gradient descent example

    - -

    -Let \( \mathbf{y} = (y_1,\cdots,y_n)^T \), \( \mathbf{\hat{y}} = (\hat{y}_1,\cdots,\hat{y}_n)^T \) and \( \beta = (\beta_0, \beta_1)^T \) - -

    -t is convenient to write \( \mathbf{\hat{y}} = X\beta \) where \( X \in \mathbb{R}^{100 \times 2} \) is the design matrix given by +

    The Hessian matrix

    +The Hessian matrix of \( C(\beta) \) is given by $$ -\begin{equation} -X \equiv \begin{bmatrix} -1 & x_1 \\ -\vdots & \vdots \\ -1 & x_{100} & \\ -\end{bmatrix}. -\tag{1} -\end{equation} +\hat{H} \equiv \begin{bmatrix} +\frac{\partial^2 C(\beta)}{\partial \beta_0^2} & \frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} \\ +\frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} & \frac{\partial^2 C(\beta)}{\partial \beta_1^2} & \\ +\end{bmatrix} = 2X^T X. $$ -The loss function is given by -$$ -C(\beta) = ||X\beta-\mathbf{y}||^2 = ||X\beta||^2 - 2 \mathbf{y}^T X\beta + ||\mathbf{y}||^2 = \sum_{i=1}^{100} (\beta_0 + \beta_1 x_i)^2 - 2 y_i (\beta_0 + \beta_1 x_i) + y_i^2 -$$ - -and we want to find \( \beta \) such that \( C(\beta) \) is minimized. +This result implies that \( C(\beta) \) is a convex function since the matrix \( X^T X \) always is positive semi-definite.

    diff --git a/doc/pub/Splines/html/._Splines-bs015.html b/doc/pub/Splines/html/._Splines-bs015.html index 1778f7a66..36d1bdfcf 100644 --- a/doc/pub/Splines/html/._Splines-bs015.html +++ b/doc/pub/Splines/html/._Splines-bs015.html @@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -128,18 +128,18 @@ MathJax.Hub.Config({

  • More on Steepest descent
  • The ideal
  • The sensitiveness of the gradient descent
  • -
  • Gradient Descent Example
  • -
  • And a corresponding example using scikit-learn
  • -
  • Convex functions
  • -
  • Convex function
  • -
  • Conditions on convex functions
  • -
  • More on convex functions
  • -
  • Some simple problems
  • -
  • Revisiting our first homework
  • -
  • Gradient descent example
  • -
  • The derivative of the cost/loss function
  • -
  • The Hessian matrix
  • -
  • Simple program
  • +
  • Convex functions
  • +
  • Convex function
  • +
  • Conditions on convex functions
  • +
  • More on convex functions
  • +
  • Some simple problems
  • +
  • Revisiting our first homework
  • +
  • Gradient descent example
  • +
  • The derivative of the cost/loss function
  • +
  • The Hessian matrix
  • +
  • Simple program
  • +
  • Gradient Descent Example
  • +
  • And a corresponding example using scikit-learn
  • Gradient descent and Ridge
  • Stochastic Gradient Descent
  • Computation of gradients
  • @@ -174,18 +174,44 @@ MathJax.Hub.Config({ -

    The derivative of the cost/loss function

    +

    Simple program

    -Computing \( \partial C(\beta) / \partial \beta_0 \) and \( \partial C(\beta) / \partial \beta_1 \) we can show that the gradient can be written as +We can now write a program that minimizes \( C(\beta) \) using the gradient descent method with a constant learning rate \( \gamma \) according to $$ -\nabla_\beta C(\beta) = (\partial C(\beta) / \partial \beta_0, \partial C(\beta) / \partial \beta_1)^T = 2\begin{bmatrix} \sum_{i=1}^{100} \left(\beta_0+\beta_1x_i-y_i\right) \\ -\sum_{i=1}^{100}\left( x_i (\beta_0+\beta_1x_i)-y_ix_i\right) \\ -\end{bmatrix} = 2X^T(X\beta - \mathbf{y}), +\beta_{k+1} = \beta_k - \gamma \nabla_\beta C(\beta_k), \ k=0,1,\cdots $$ -where \( X \) is the design matrix defined above. +

    +We can use the expression we computed for the gradient and let use a +\( \beta_0 \) be chosen randomly and let \( \gamma = 0.001 \). Stop iterating +when \( ||\nabla_\beta C(\beta_k) || \leq \epsilon = 10^{-8} \). +

    +And finally we can compare our solution for \( \beta \) with the analytic result given by +\( \beta= (X^TX)^{-1} X^T \mathbf{y} \). +

    + + +

    import numpy as np
    +
    +"""
    +The following setup is just a suggestion, feel free to write it the way you like.
    +"""
    +
    +#Setup problem described in the exercise
    +N  = 100 #Nr of datapoints
    +M  = 2 #Nr of features
    +x  = np.random.rand(N) #Uniformly generated x-values in [0,1]
    +y  = 5*x**2 + 0.1*np.random.randn(N)
    +X  = np.c_[np.ones(N),x] #Construct design matrix
    +
    +#Compute beta according to normal equations to compare with GD solution
    +Xt_X_inv = np.linalg.inv(np.dot(X.T,X))
    +Xt_y     = np.dot(X.transpose(),y)
    +beta_NE = np.dot(Xt_X_inv,Xt_y)
    +print(beta_NE)
    +

    diff --git a/doc/pub/Splines/html/._Splines-bs016.html b/doc/pub/Splines/html/._Splines-bs016.html index d4a0a6cb0..8f6c8a419 100644 --- a/doc/pub/Splines/html/._Splines-bs016.html +++ b/doc/pub/Splines/html/._Splines-bs016.html @@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -128,18 +128,18 @@ MathJax.Hub.Config({

  • More on Steepest descent
  • The ideal
  • The sensitiveness of the gradient descent
  • -
  • Gradient Descent Example
  • -
  • And a corresponding example using scikit-learn
  • -
  • Convex functions
  • -
  • Convex function
  • -
  • Conditions on convex functions
  • -
  • More on convex functions
  • -
  • Some simple problems
  • -
  • Revisiting our first homework
  • -
  • Gradient descent example
  • -
  • The derivative of the cost/loss function
  • -
  • The Hessian matrix
  • -
  • Simple program
  • +
  • Convex functions
  • +
  • Convex function
  • +
  • Conditions on convex functions
  • +
  • More on convex functions
  • +
  • Some simple problems
  • +
  • Revisiting our first homework
  • +
  • Gradient descent example
  • +
  • The derivative of the cost/loss function
  • +
  • The Hessian matrix
  • +
  • Simple program
  • +
  • Gradient Descent Example
  • +
  • And a corresponding example using scikit-learn
  • Gradient descent and Ridge
  • Stochastic Gradient Descent
  • Computation of gradients
  • @@ -174,17 +174,52 @@ MathJax.Hub.Config({ -

    The Hessian matrix

    -The Hessian matrix of \( C(\beta) \) is given by -$$ -\hat{H} \equiv \begin{bmatrix} -\frac{\partial^2 C(\beta)}{\partial \beta_0^2} & \frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} \\ -\frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} & \frac{\partial^2 C(\beta)}{\partial \beta_1^2} & \\ -\end{bmatrix} = 2X^T X. -$$ +

    Gradient Descent Example

    -This result implies that \( C(\beta) \) is a convex function since the matrix \( X^T X \) always is positive semi-definite. +

    +Another simple example is here +

    + +

    # Importing various packages
    +from random import random, seed
    +import numpy as np
    +import matplotlib.pyplot as plt
    +from mpl_toolkits.mplot3d import Axes3D
    +from matplotlib import cm
    +from matplotlib.ticker import LinearLocator, FormatStrFormatter
    +import sys
    +
    +x = 2*np.random.rand(100,1)
    +y = 4+3*x+np.random.randn(100,1)
    +
    +xb = np.c_[np.ones((100,1)), x]
    +beta_linreg = np.linalg.inv(xb.T.dot(xb)).dot(xb.T).dot(y)
    +print(beta_linreg)
    +beta = np.random.randn(2,1)
    +
    +eta = 0.1
    +Niterations = 1000
    +m = 100
    +
    +for iter in range(Niterations):
    +    gradients = 2.0/m*xb.T.dot(xb.dot(beta)-y)
    +    beta -= eta*gradients
    +
    +print(beta)
    +xnew = np.array([[0],[2]])
    +xbnew = np.c_[np.ones((2,1)), xnew]
    +ypredict = xbnew.dot(beta)
    +ypredict2 = xbnew.dot(beta_linreg)
    +plt.plot(xnew, ypredict, "r-")
    +plt.plot(xnew, ypredict2, "b-")
    +plt.plot(x, y ,'ro')
    +plt.axis([0,2.0,0, 15.0])
    +plt.xlabel(r'$x$')
    +plt.ylabel(r'$y$')
    +plt.title(r'Gradient descent example')
    +plt.show()
    +

    diff --git a/doc/pub/Splines/html/._Splines-bs017.html b/doc/pub/Splines/html/._Splines-bs017.html index d02aaae45..b6e338b6b 100644 --- a/doc/pub/Splines/html/._Splines-bs017.html +++ b/doc/pub/Splines/html/._Splines-bs017.html @@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -128,18 +128,18 @@ MathJax.Hub.Config({

  • More on Steepest descent
  • The ideal
  • The sensitiveness of the gradient descent
  • -
  • Gradient Descent Example
  • -
  • And a corresponding example using scikit-learn
  • -
  • Convex functions
  • -
  • Convex function
  • -
  • Conditions on convex functions
  • -
  • More on convex functions
  • -
  • Some simple problems
  • -
  • Revisiting our first homework
  • -
  • Gradient descent example
  • -
  • The derivative of the cost/loss function
  • -
  • The Hessian matrix
  • -
  • Simple program
  • +
  • Convex functions
  • +
  • Convex function
  • +
  • Conditions on convex functions
  • +
  • More on convex functions
  • +
  • Some simple problems
  • +
  • Revisiting our first homework
  • +
  • Gradient descent example
  • +
  • The derivative of the cost/loss function
  • +
  • The Hessian matrix
  • +
  • Simple program
  • +
  • Gradient Descent Example
  • +
  • And a corresponding example using scikit-learn
  • Gradient descent and Ridge
  • Stochastic Gradient Descent
  • Computation of gradients
  • @@ -174,43 +174,26 @@ MathJax.Hub.Config({ -

    Simple program

    +

    And a corresponding example using scikit-learn

    -

    -We can now write a program that minimizes \( C(\beta) \) using the gradient descent method with a constant learning rate \( \gamma \) according to -$$ -\beta_{k+1} = \beta_k - \gamma \nabla_\beta C(\beta_k), \ k=0,1,\cdots -$$ - -

    -We can use the expression we computed for the gradient and let use a -\( \beta_0 \) be chosen randomly and let \( \gamma = 0.001 \). Stop iterating -when \( ||\nabla_\beta C(\beta_k) || < \epsilon = 10^{-8} \). - -

    -And finally we can compare our solution for \( \beta \) with the analytic result given by -\( \beta= (X^TX)^{-1} X^T \mathbf{y} \).

    -

    import numpy as np
    +
    # Importing various packages
    +from random import random, seed
    +import numpy as np
    +import matplotlib.pyplot as plt
    +from sklearn.linear_model import SGDRegressor
     
    -"""
    -The following setup is just a suggestion, feel free to write it the way you like.
    -"""
    +x = 2*np.random.rand(100,1)
    +y = 4+3*x+np.random.randn(100,1)
     
    -#Setup problem described in the exercise
    -N  = 100 #Nr of datapoints
    -M  = 2 #Nr of features
    -x  = np.random.rand(N) #Uniformly generated x-values in [0,1]
    -y  = 5*x**2 + 0.1*np.random.randn(N)
    -X  = np.c_[np.ones(N),x] #Construct design matrix
    -
    -#Compute beta according to normal equations to compare with GD solution
    -Xt_X_inv = np.linalg.inv(np.dot(X.T,X))
    -Xt_y     = np.dot(X.transpose(),y)
    -beta_NE = np.dot(Xt_X_inv,Xt_y)
    -print(beta_NE)
    +xb = np.c_[np.ones((100,1)), x]
    +beta_linreg = np.linalg.inv(xb.T.dot(xb)).dot(xb.T).dot(y)
    +print(beta_linreg)
    +sgdreg = SGDRegressor(n_iter = 50, penalty=None, eta0=0.1)
    +sgdreg.fit(x,y.ravel())
    +print(sgdreg.intercept_, sgdreg.coef_)
     

    diff --git a/doc/pub/Splines/html/._Splines-bs018.html b/doc/pub/Splines/html/._Splines-bs018.html index 878874987..507b4a0f3 100644 --- a/doc/pub/Splines/html/._Splines-bs018.html +++ b/doc/pub/Splines/html/._Splines-bs018.html @@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -128,18 +128,18 @@ MathJax.Hub.Config({

  • More on Steepest descent
  • The ideal
  • The sensitiveness of the gradient descent
  • -
  • Gradient Descent Example
  • -
  • And a corresponding example using scikit-learn
  • -
  • Convex functions
  • -
  • Convex function
  • -
  • Conditions on convex functions
  • -
  • More on convex functions
  • -
  • Some simple problems
  • -
  • Revisiting our first homework
  • -
  • Gradient descent example
  • -
  • The derivative of the cost/loss function
  • -
  • The Hessian matrix
  • -
  • Simple program
  • +
  • Convex functions
  • +
  • Convex function
  • +
  • Conditions on convex functions
  • +
  • More on convex functions
  • +
  • Some simple problems
  • +
  • Revisiting our first homework
  • +
  • Gradient descent example
  • +
  • The derivative of the cost/loss function
  • +
  • The Hessian matrix
  • +
  • Simple program
  • +
  • Gradient Descent Example
  • +
  • And a corresponding example using scikit-learn
  • Gradient descent and Ridge
  • Stochastic Gradient Descent
  • Computation of gradients
  • diff --git a/doc/pub/Splines/html/._Splines-bs019.html b/doc/pub/Splines/html/._Splines-bs019.html index 464a234f4..dda06818b 100644 --- a/doc/pub/Splines/html/._Splines-bs019.html +++ b/doc/pub/Splines/html/._Splines-bs019.html @@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -128,18 +128,18 @@ MathJax.Hub.Config({
  • More on Steepest descent
  • The ideal
  • The sensitiveness of the gradient descent
  • -
  • Gradient Descent Example
  • -
  • And a corresponding example using scikit-learn
  • -
  • Convex functions
  • -
  • Convex function
  • -
  • Conditions on convex functions
  • -
  • More on convex functions
  • -
  • Some simple problems
  • -
  • Revisiting our first homework
  • -
  • Gradient descent example
  • -
  • The derivative of the cost/loss function
  • -
  • The Hessian matrix
  • -
  • Simple program
  • +
  • Convex functions
  • +
  • Convex function
  • +
  • Conditions on convex functions
  • +
  • More on convex functions
  • +
  • Some simple problems
  • +
  • Revisiting our first homework
  • +
  • Gradient descent example
  • +
  • The derivative of the cost/loss function
  • +
  • The Hessian matrix
  • +
  • Simple program
  • +
  • Gradient Descent Example
  • +
  • And a corresponding example using scikit-learn
  • Gradient descent and Ridge
  • Stochastic Gradient Descent
  • Computation of gradients
  • @@ -183,7 +183,7 @@ the shortcomings of the Gradient descent method discussed above.

    The underlying idea of SGD comes from the observation that the cost function, which we want to minimize, can almost always be written as a -sum over \( n \) datapoints \( \{\mathbf{x}_i\}_{i=1}^n \), +sum over \( n \) data points \( \{\mathbf{x}_i\}_{i=1}^n \), $$ C(\mathbf{\beta}) = \sum_{i=1}^n c_i(\mathbf{x}_i, \mathbf{\beta}). diff --git a/doc/pub/Splines/html/._Splines-bs020.html b/doc/pub/Splines/html/._Splines-bs020.html index 953411294..2b13ebb75 100644 --- a/doc/pub/Splines/html/._Splines-bs020.html +++ b/doc/pub/Splines/html/._Splines-bs020.html @@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -128,18 +128,18 @@ MathJax.Hub.Config({

  • More on Steepest descent
  • The ideal
  • The sensitiveness of the gradient descent
  • -
  • Gradient Descent Example
  • -
  • And a corresponding example using scikit-learn
  • -
  • Convex functions
  • -
  • Convex function
  • -
  • Conditions on convex functions
  • -
  • More on convex functions
  • -
  • Some simple problems
  • -
  • Revisiting our first homework
  • -
  • Gradient descent example
  • -
  • The derivative of the cost/loss function
  • -
  • The Hessian matrix
  • -
  • Simple program
  • +
  • Convex functions
  • +
  • Convex function
  • +
  • Conditions on convex functions
  • +
  • More on convex functions
  • +
  • Some simple problems
  • +
  • Revisiting our first homework
  • +
  • Gradient descent example
  • +
  • The derivative of the cost/loss function
  • +
  • The Hessian matrix
  • +
  • Simple program
  • +
  • Gradient Descent Example
  • +
  • And a corresponding example using scikit-learn
  • Gradient descent and Ridge
  • Stochastic Gradient Descent
  • Computation of gradients
  • @@ -187,7 +187,7 @@ $$

    Stochasticity/randomness is introduced by only taking the gradient on a subset of the data called minibatches. If there are \( n \) -datapoints and the size of each minibatch is \( M \), there will be \( n/M \) +data points and the size of each minibatch is \( M \), there will be \( n/M \) minibatches. We denote these minibatches by \( B_k \) where \( k=1,\cdots,n/M \). diff --git a/doc/pub/Splines/html/._Splines-bs021.html b/doc/pub/Splines/html/._Splines-bs021.html index 062ea2901..6525dab00 100644 --- a/doc/pub/Splines/html/._Splines-bs021.html +++ b/doc/pub/Splines/html/._Splines-bs021.html @@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -128,18 +128,18 @@ MathJax.Hub.Config({

  • More on Steepest descent
  • The ideal
  • The sensitiveness of the gradient descent
  • -
  • Gradient Descent Example
  • -
  • And a corresponding example using scikit-learn
  • -
  • Convex functions
  • -
  • Convex function
  • -
  • Conditions on convex functions
  • -
  • More on convex functions
  • -
  • Some simple problems
  • -
  • Revisiting our first homework
  • -
  • Gradient descent example
  • -
  • The derivative of the cost/loss function
  • -
  • The Hessian matrix
  • -
  • Simple program
  • +
  • Convex functions
  • +
  • Convex function
  • +
  • Conditions on convex functions
  • +
  • More on convex functions
  • +
  • Some simple problems
  • +
  • Revisiting our first homework
  • +
  • Gradient descent example
  • +
  • The derivative of the cost/loss function
  • +
  • The Hessian matrix
  • +
  • Simple program
  • +
  • Gradient Descent Example
  • +
  • And a corresponding example using scikit-learn
  • Gradient descent and Ridge
  • Stochastic Gradient Descent
  • Computation of gradients
  • @@ -175,21 +175,21 @@ MathJax.Hub.Config({

    SGD example

    -As an example, suppose we have \( 10 \) datapoints \( ( \mathbf{x}_1, -\cdots, \mathbf{x}_{10} ) \) and we choose to have \( M=5 \) minibathces, -then each minibatch contains two datapoints. In particular we have +As an example, suppose we have \( 10 \) data points \( (\mathbf{x}_1,\cdots, \mathbf{x}_{10}) \) +and we choose to have \( M=5 \) minibathces, +then each minibatch contains two data points. In particular we have \( B_1 = (\mathbf{x}_1,\mathbf{x}_2), \cdots, B_5 = (\mathbf{x}_9,\mathbf{x}_{10}) \). Note that if you choose \( M=1 \) you -have only a single batch with all datapoints and on the other extreme, +have only a single batch with all data points and on the other extreme, you may choose \( M=n \) resulting in a minibatch for each datapoint, i.e \( B_k = \mathbf{x}_k \).

    The idea is now to approximate the gradient by replacing the sum over -all datapoints with a sum over the datapoints in one the minibatches +all data points with a sum over the data points in one the minibatches picked at random in each gradient descent step $$ -\nabla_\beta +\nabla_{\beta} C(\mathbf{\beta}) = \sum_{i=1}^n \nabla_\beta c_i(\mathbf{x}_i, \mathbf{\beta}) \rightarrow \sum_{i \in B_k}^n \nabla_\beta c_i(\mathbf{x}_i, \mathbf{\beta}). diff --git a/doc/pub/Splines/html/._Splines-bs022.html b/doc/pub/Splines/html/._Splines-bs022.html index 6dc924616..c7901fea0 100644 --- a/doc/pub/Splines/html/._Splines-bs022.html +++ b/doc/pub/Splines/html/._Splines-bs022.html @@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -128,18 +128,18 @@ MathJax.Hub.Config({

  • More on Steepest descent
  • The ideal
  • The sensitiveness of the gradient descent
  • -
  • Gradient Descent Example
  • -
  • And a corresponding example using scikit-learn
  • -
  • Convex functions
  • -
  • Convex function
  • -
  • Conditions on convex functions
  • -
  • More on convex functions
  • -
  • Some simple problems
  • -
  • Revisiting our first homework
  • -
  • Gradient descent example
  • -
  • The derivative of the cost/loss function
  • -
  • The Hessian matrix
  • -
  • Simple program
  • +
  • Convex functions
  • +
  • Convex function
  • +
  • Conditions on convex functions
  • +
  • More on convex functions
  • +
  • Some simple problems
  • +
  • Revisiting our first homework
  • +
  • Gradient descent example
  • +
  • The derivative of the cost/loss function
  • +
  • The Hessian matrix
  • +
  • Simple program
  • +
  • Gradient Descent Example
  • +
  • And a corresponding example using scikit-learn
  • Gradient descent and Ridge
  • Stochastic Gradient Descent
  • Computation of gradients
  • diff --git a/doc/pub/Splines/html/._Splines-bs023.html b/doc/pub/Splines/html/._Splines-bs023.html index 9c98ea73d..ccc0d0fab 100644 --- a/doc/pub/Splines/html/._Splines-bs023.html +++ b/doc/pub/Splines/html/._Splines-bs023.html @@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -128,18 +128,18 @@ MathJax.Hub.Config({
  • More on Steepest descent
  • The ideal
  • The sensitiveness of the gradient descent
  • -
  • Gradient Descent Example
  • -
  • And a corresponding example using scikit-learn
  • -
  • Convex functions
  • -
  • Convex function
  • -
  • Conditions on convex functions
  • -
  • More on convex functions
  • -
  • Some simple problems
  • -
  • Revisiting our first homework
  • -
  • Gradient descent example
  • -
  • The derivative of the cost/loss function
  • -
  • The Hessian matrix
  • -
  • Simple program
  • +
  • Convex functions
  • +
  • Convex function
  • +
  • Conditions on convex functions
  • +
  • More on convex functions
  • +
  • Some simple problems
  • +
  • Revisiting our first homework
  • +
  • Gradient descent example
  • +
  • The derivative of the cost/loss function
  • +
  • The Hessian matrix
  • +
  • Simple program
  • +
  • Gradient Descent Example
  • +
  • And a corresponding example using scikit-learn
  • Gradient descent and Ridge
  • Stochastic Gradient Descent
  • Computation of gradients
  • @@ -199,8 +199,8 @@ Taking the gradient only on a subset of the data has two important benefits. First, it introduces randomness which decreases the chance that our opmization scheme gets stuck in a local minima. Second, if the size of the minibatches are small relative to the number of -datapoints (\( M < n \)), the computation of the gradient is much -cheaper since we sum over the datapoints in the k-th minibatch and not +datapoints (\( M < n \)), the computation of the gradient is much +cheaper since we sum over the datapoints in the \( k-th \) minibatch and not all \( n \) datapoints.

    diff --git a/doc/pub/Splines/html/._Splines-bs024.html b/doc/pub/Splines/html/._Splines-bs024.html index 1653551d2..084a9a24f 100644 --- a/doc/pub/Splines/html/._Splines-bs024.html +++ b/doc/pub/Splines/html/._Splines-bs024.html @@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -128,18 +128,18 @@ MathJax.Hub.Config({

  • More on Steepest descent
  • The ideal
  • The sensitiveness of the gradient descent
  • -
  • Gradient Descent Example
  • -
  • And a corresponding example using scikit-learn
  • -
  • Convex functions
  • -
  • Convex function
  • -
  • Conditions on convex functions
  • -
  • More on convex functions
  • -
  • Some simple problems
  • -
  • Revisiting our first homework
  • -
  • Gradient descent example
  • -
  • The derivative of the cost/loss function
  • -
  • The Hessian matrix
  • -
  • Simple program
  • +
  • Convex functions
  • +
  • Convex function
  • +
  • Conditions on convex functions
  • +
  • More on convex functions
  • +
  • Some simple problems
  • +
  • Revisiting our first homework
  • +
  • Gradient descent example
  • +
  • The derivative of the cost/loss function
  • +
  • The Hessian matrix
  • +
  • Simple program
  • +
  • Gradient Descent Example
  • +
  • And a corresponding example using scikit-learn
  • Gradient descent and Ridge
  • Stochastic Gradient Descent
  • Computation of gradients
  • diff --git a/doc/pub/Splines/html/._Splines-bs025.html b/doc/pub/Splines/html/._Splines-bs025.html index 49c8b196a..2568a5049 100644 --- a/doc/pub/Splines/html/._Splines-bs025.html +++ b/doc/pub/Splines/html/._Splines-bs025.html @@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -128,18 +128,18 @@ MathJax.Hub.Config({
  • More on Steepest descent
  • The ideal
  • The sensitiveness of the gradient descent
  • -
  • Gradient Descent Example
  • -
  • And a corresponding example using scikit-learn
  • -
  • Convex functions
  • -
  • Convex function
  • -
  • Conditions on convex functions
  • -
  • More on convex functions
  • -
  • Some simple problems
  • -
  • Revisiting our first homework
  • -
  • Gradient descent example
  • -
  • The derivative of the cost/loss function
  • -
  • The Hessian matrix
  • -
  • Simple program
  • +
  • Convex functions
  • +
  • Convex function
  • +
  • Conditions on convex functions
  • +
  • More on convex functions
  • +
  • Some simple problems
  • +
  • Revisiting our first homework
  • +
  • Gradient descent example
  • +
  • The derivative of the cost/loss function
  • +
  • The Hessian matrix
  • +
  • Simple program
  • +
  • Gradient Descent Example
  • +
  • And a corresponding example using scikit-learn
  • Gradient descent and Ridge
  • Stochastic Gradient Descent
  • Computation of gradients
  • @@ -182,7 +182,7 @@ number of epochs in such a way that it becomes very small after a reasonable time such that we do not move at all.

    -As an example, let \( e = 0,1,2,3,\cdots \) denote the current epoch and let \( t_0, t_1 > 0 \) be two fixed numbers. Furthermore, let \( t = e \cdot m + i \) where \( m \) is the number of minibatches and \( i=0,\cdots,m-1 \). Then the function $$\gamma_j(t; t_0, t_1) = \frac{t_0}{t+t_1} $$ goes to zero as the number of epochs gets large. I.e. we start with a step length \( \gamma_j (0; t_0, t_1) = t_0/t_1 \) which decays in time \( t \). +As an example, let \( e = 0,1,2,3,\cdots \) denote the current epoch and let \( t_0, t_1 > 0 \) be two fixed numbers. Furthermore, let \( t = e \cdot m + i \) where \( m \) is the number of minibatches and \( i=0,\cdots,m-1 \). Then the function $$\gamma_j(t; t_0, t_1) = \frac{t_0}{t+t_1} $$ goes to zero as the number of epochs gets large. I.e. we start with a step length \( \gamma_j (0; t_0, t_1) = t_0/t_1 \) which decays in time \( t \).

    In this way we can fix the number of epochs, compute \( \beta \) and diff --git a/doc/pub/Splines/html/._Splines-bs026.html b/doc/pub/Splines/html/._Splines-bs026.html index a3bc73c10..069443c14 100644 --- a/doc/pub/Splines/html/._Splines-bs026.html +++ b/doc/pub/Splines/html/._Splines-bs026.html @@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -128,18 +128,18 @@ MathJax.Hub.Config({

  • More on Steepest descent
  • The ideal
  • The sensitiveness of the gradient descent
  • -
  • Gradient Descent Example
  • -
  • And a corresponding example using scikit-learn
  • -
  • Convex functions
  • -
  • Convex function
  • -
  • Conditions on convex functions
  • -
  • More on convex functions
  • -
  • Some simple problems
  • -
  • Revisiting our first homework
  • -
  • Gradient descent example
  • -
  • The derivative of the cost/loss function
  • -
  • The Hessian matrix
  • -
  • Simple program
  • +
  • Convex functions
  • +
  • Convex function
  • +
  • Conditions on convex functions
  • +
  • More on convex functions
  • +
  • Some simple problems
  • +
  • Revisiting our first homework
  • +
  • Gradient descent example
  • +
  • The derivative of the cost/loss function
  • +
  • The Hessian matrix
  • +
  • Simple program
  • +
  • Gradient Descent Example
  • +
  • And a corresponding example using scikit-learn
  • Gradient descent and Ridge
  • Stochastic Gradient Descent
  • Computation of gradients
  • diff --git a/doc/pub/Splines/html/._Splines-bs027.html b/doc/pub/Splines/html/._Splines-bs027.html index aa6377c6d..faf2c87d5 100644 --- a/doc/pub/Splines/html/._Splines-bs027.html +++ b/doc/pub/Splines/html/._Splines-bs027.html @@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -128,18 +128,18 @@ MathJax.Hub.Config({
  • More on Steepest descent
  • The ideal
  • The sensitiveness of the gradient descent
  • -
  • Gradient Descent Example
  • -
  • And a corresponding example using scikit-learn
  • -
  • Convex functions
  • -
  • Convex function
  • -
  • Conditions on convex functions
  • -
  • More on convex functions
  • -
  • Some simple problems
  • -
  • Revisiting our first homework
  • -
  • Gradient descent example
  • -
  • The derivative of the cost/loss function
  • -
  • The Hessian matrix
  • -
  • Simple program
  • +
  • Convex functions
  • +
  • Convex function
  • +
  • Conditions on convex functions
  • +
  • More on convex functions
  • +
  • Some simple problems
  • +
  • Revisiting our first homework
  • +
  • Gradient descent example
  • +
  • The derivative of the cost/loss function
  • +
  • The Hessian matrix
  • +
  • Simple program
  • +
  • Gradient Descent Example
  • +
  • And a corresponding example using scikit-learn
  • Gradient descent and Ridge
  • Stochastic Gradient Descent
  • Computation of gradients
  • diff --git a/doc/pub/Splines/html/._Splines-bs028.html b/doc/pub/Splines/html/._Splines-bs028.html index 07849d947..b6e89d1b7 100644 --- a/doc/pub/Splines/html/._Splines-bs028.html +++ b/doc/pub/Splines/html/._Splines-bs028.html @@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -128,18 +128,18 @@ MathJax.Hub.Config({
  • More on Steepest descent
  • The ideal
  • The sensitiveness of the gradient descent
  • -
  • Gradient Descent Example
  • -
  • And a corresponding example using scikit-learn
  • -
  • Convex functions
  • -
  • Convex function
  • -
  • Conditions on convex functions
  • -
  • More on convex functions
  • -
  • Some simple problems
  • -
  • Revisiting our first homework
  • -
  • Gradient descent example
  • -
  • The derivative of the cost/loss function
  • -
  • The Hessian matrix
  • -
  • Simple program
  • +
  • Convex functions
  • +
  • Convex function
  • +
  • Conditions on convex functions
  • +
  • More on convex functions
  • +
  • Some simple problems
  • +
  • Revisiting our first homework
  • +
  • Gradient descent example
  • +
  • The derivative of the cost/loss function
  • +
  • The Hessian matrix
  • +
  • Simple program
  • +
  • Gradient Descent Example
  • +
  • And a corresponding example using scikit-learn
  • Gradient descent and Ridge
  • Stochastic Gradient Descent
  • Computation of gradients
  • diff --git a/doc/pub/Splines/html/._Splines-bs029.html b/doc/pub/Splines/html/._Splines-bs029.html index acbc6471d..7ea3b65e7 100644 --- a/doc/pub/Splines/html/._Splines-bs029.html +++ b/doc/pub/Splines/html/._Splines-bs029.html @@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -128,18 +128,18 @@ MathJax.Hub.Config({
  • More on Steepest descent
  • The ideal
  • The sensitiveness of the gradient descent
  • -
  • Gradient Descent Example
  • -
  • And a corresponding example using scikit-learn
  • -
  • Convex functions
  • -
  • Convex function
  • -
  • Conditions on convex functions
  • -
  • More on convex functions
  • -
  • Some simple problems
  • -
  • Revisiting our first homework
  • -
  • Gradient descent example
  • -
  • The derivative of the cost/loss function
  • -
  • The Hessian matrix
  • -
  • Simple program
  • +
  • Convex functions
  • +
  • Convex function
  • +
  • Conditions on convex functions
  • +
  • More on convex functions
  • +
  • Some simple problems
  • +
  • Revisiting our first homework
  • +
  • Gradient descent example
  • +
  • The derivative of the cost/loss function
  • +
  • The Hessian matrix
  • +
  • Simple program
  • +
  • Gradient Descent Example
  • +
  • And a corresponding example using scikit-learn
  • Gradient descent and Ridge
  • Stochastic Gradient Descent
  • Computation of gradients
  • diff --git a/doc/pub/Splines/html/._Splines-bs030.html b/doc/pub/Splines/html/._Splines-bs030.html index 2ffe81cff..03448fbfc 100644 --- a/doc/pub/Splines/html/._Splines-bs030.html +++ b/doc/pub/Splines/html/._Splines-bs030.html @@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -128,18 +128,18 @@ MathJax.Hub.Config({
  • More on Steepest descent
  • The ideal
  • The sensitiveness of the gradient descent
  • -
  • Gradient Descent Example
  • -
  • And a corresponding example using scikit-learn
  • -
  • Convex functions
  • -
  • Convex function
  • -
  • Conditions on convex functions
  • -
  • More on convex functions
  • -
  • Some simple problems
  • -
  • Revisiting our first homework
  • -
  • Gradient descent example
  • -
  • The derivative of the cost/loss function
  • -
  • The Hessian matrix
  • -
  • Simple program
  • +
  • Convex functions
  • +
  • Convex function
  • +
  • Conditions on convex functions
  • +
  • More on convex functions
  • +
  • Some simple problems
  • +
  • Revisiting our first homework
  • +
  • Gradient descent example
  • +
  • The derivative of the cost/loss function
  • +
  • The Hessian matrix
  • +
  • Simple program
  • +
  • Gradient Descent Example
  • +
  • And a corresponding example using scikit-learn
  • Gradient descent and Ridge
  • Stochastic Gradient Descent
  • Computation of gradients
  • diff --git a/doc/pub/Splines/html/._Splines-bs031.html b/doc/pub/Splines/html/._Splines-bs031.html index ad740e7f8..0237a96c7 100644 --- a/doc/pub/Splines/html/._Splines-bs031.html +++ b/doc/pub/Splines/html/._Splines-bs031.html @@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -128,18 +128,18 @@ MathJax.Hub.Config({
  • More on Steepest descent
  • The ideal
  • The sensitiveness of the gradient descent
  • -
  • Gradient Descent Example
  • -
  • And a corresponding example using scikit-learn
  • -
  • Convex functions
  • -
  • Convex function
  • -
  • Conditions on convex functions
  • -
  • More on convex functions
  • -
  • Some simple problems
  • -
  • Revisiting our first homework
  • -
  • Gradient descent example
  • -
  • The derivative of the cost/loss function
  • -
  • The Hessian matrix
  • -
  • Simple program
  • +
  • Convex functions
  • +
  • Convex function
  • +
  • Conditions on convex functions
  • +
  • More on convex functions
  • +
  • Some simple problems
  • +
  • Revisiting our first homework
  • +
  • Gradient descent example
  • +
  • The derivative of the cost/loss function
  • +
  • The Hessian matrix
  • +
  • Simple program
  • +
  • Gradient Descent Example
  • +
  • And a corresponding example using scikit-learn
  • Gradient descent and Ridge
  • Stochastic Gradient Descent
  • Computation of gradients
  • diff --git a/doc/pub/Splines/html/._Splines-bs032.html b/doc/pub/Splines/html/._Splines-bs032.html index a4ecf522e..0c101dce5 100644 --- a/doc/pub/Splines/html/._Splines-bs032.html +++ b/doc/pub/Splines/html/._Splines-bs032.html @@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -128,18 +128,18 @@ MathJax.Hub.Config({
  • More on Steepest descent
  • The ideal
  • The sensitiveness of the gradient descent
  • -
  • Gradient Descent Example
  • -
  • And a corresponding example using scikit-learn
  • -
  • Convex functions
  • -
  • Convex function
  • -
  • Conditions on convex functions
  • -
  • More on convex functions
  • -
  • Some simple problems
  • -
  • Revisiting our first homework
  • -
  • Gradient descent example
  • -
  • The derivative of the cost/loss function
  • -
  • The Hessian matrix
  • -
  • Simple program
  • +
  • Convex functions
  • +
  • Convex function
  • +
  • Conditions on convex functions
  • +
  • More on convex functions
  • +
  • Some simple problems
  • +
  • Revisiting our first homework
  • +
  • Gradient descent example
  • +
  • The derivative of the cost/loss function
  • +
  • The Hessian matrix
  • +
  • Simple program
  • +
  • Gradient Descent Example
  • +
  • And a corresponding example using scikit-learn
  • Gradient descent and Ridge
  • Stochastic Gradient Descent
  • Computation of gradients
  • diff --git a/doc/pub/Splines/html/._Splines-bs033.html b/doc/pub/Splines/html/._Splines-bs033.html index 5ca815020..ea7e4344b 100644 --- a/doc/pub/Splines/html/._Splines-bs033.html +++ b/doc/pub/Splines/html/._Splines-bs033.html @@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -128,18 +128,18 @@ MathJax.Hub.Config({
  • More on Steepest descent
  • The ideal
  • The sensitiveness of the gradient descent
  • -
  • Gradient Descent Example
  • -
  • And a corresponding example using scikit-learn
  • -
  • Convex functions
  • -
  • Convex function
  • -
  • Conditions on convex functions
  • -
  • More on convex functions
  • -
  • Some simple problems
  • -
  • Revisiting our first homework
  • -
  • Gradient descent example
  • -
  • The derivative of the cost/loss function
  • -
  • The Hessian matrix
  • -
  • Simple program
  • +
  • Convex functions
  • +
  • Convex function
  • +
  • Conditions on convex functions
  • +
  • More on convex functions
  • +
  • Some simple problems
  • +
  • Revisiting our first homework
  • +
  • Gradient descent example
  • +
  • The derivative of the cost/loss function
  • +
  • The Hessian matrix
  • +
  • Simple program
  • +
  • Gradient Descent Example
  • +
  • And a corresponding example using scikit-learn
  • Gradient descent and Ridge
  • Stochastic Gradient Descent
  • Computation of gradients
  • diff --git a/doc/pub/Splines/html/._Splines-bs034.html b/doc/pub/Splines/html/._Splines-bs034.html index 33e8bbc77..f96293d33 100644 --- a/doc/pub/Splines/html/._Splines-bs034.html +++ b/doc/pub/Splines/html/._Splines-bs034.html @@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -128,18 +128,18 @@ MathJax.Hub.Config({
  • More on Steepest descent
  • The ideal
  • The sensitiveness of the gradient descent
  • -
  • Gradient Descent Example
  • -
  • And a corresponding example using scikit-learn
  • -
  • Convex functions
  • -
  • Convex function
  • -
  • Conditions on convex functions
  • -
  • More on convex functions
  • -
  • Some simple problems
  • -
  • Revisiting our first homework
  • -
  • Gradient descent example
  • -
  • The derivative of the cost/loss function
  • -
  • The Hessian matrix
  • -
  • Simple program
  • +
  • Convex functions
  • +
  • Convex function
  • +
  • Conditions on convex functions
  • +
  • More on convex functions
  • +
  • Some simple problems
  • +
  • Revisiting our first homework
  • +
  • Gradient descent example
  • +
  • The derivative of the cost/loss function
  • +
  • The Hessian matrix
  • +
  • Simple program
  • +
  • Gradient Descent Example
  • +
  • And a corresponding example using scikit-learn
  • Gradient descent and Ridge
  • Stochastic Gradient Descent
  • Computation of gradients
  • diff --git a/doc/pub/Splines/html/._Splines-bs035.html b/doc/pub/Splines/html/._Splines-bs035.html index a783f024b..63617ce8f 100644 --- a/doc/pub/Splines/html/._Splines-bs035.html +++ b/doc/pub/Splines/html/._Splines-bs035.html @@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -128,18 +128,18 @@ MathJax.Hub.Config({
  • More on Steepest descent
  • The ideal
  • The sensitiveness of the gradient descent
  • -
  • Gradient Descent Example
  • -
  • And a corresponding example using scikit-learn
  • -
  • Convex functions
  • -
  • Convex function
  • -
  • Conditions on convex functions
  • -
  • More on convex functions
  • -
  • Some simple problems
  • -
  • Revisiting our first homework
  • -
  • Gradient descent example
  • -
  • The derivative of the cost/loss function
  • -
  • The Hessian matrix
  • -
  • Simple program
  • +
  • Convex functions
  • +
  • Convex function
  • +
  • Conditions on convex functions
  • +
  • More on convex functions
  • +
  • Some simple problems
  • +
  • Revisiting our first homework
  • +
  • Gradient descent example
  • +
  • The derivative of the cost/loss function
  • +
  • The Hessian matrix
  • +
  • Simple program
  • +
  • Gradient Descent Example
  • +
  • And a corresponding example using scikit-learn
  • Gradient descent and Ridge
  • Stochastic Gradient Descent
  • Computation of gradients
  • diff --git a/doc/pub/Splines/html/._Splines-bs036.html b/doc/pub/Splines/html/._Splines-bs036.html index 4070ced3b..93e47979b 100644 --- a/doc/pub/Splines/html/._Splines-bs036.html +++ b/doc/pub/Splines/html/._Splines-bs036.html @@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -128,18 +128,18 @@ MathJax.Hub.Config({
  • More on Steepest descent
  • The ideal
  • The sensitiveness of the gradient descent
  • -
  • Gradient Descent Example
  • -
  • And a corresponding example using scikit-learn
  • -
  • Convex functions
  • -
  • Convex function
  • -
  • Conditions on convex functions
  • -
  • More on convex functions
  • -
  • Some simple problems
  • -
  • Revisiting our first homework
  • -
  • Gradient descent example
  • -
  • The derivative of the cost/loss function
  • -
  • The Hessian matrix
  • -
  • Simple program
  • +
  • Convex functions
  • +
  • Convex function
  • +
  • Conditions on convex functions
  • +
  • More on convex functions
  • +
  • Some simple problems
  • +
  • Revisiting our first homework
  • +
  • Gradient descent example
  • +
  • The derivative of the cost/loss function
  • +
  • The Hessian matrix
  • +
  • Simple program
  • +
  • Gradient Descent Example
  • +
  • And a corresponding example using scikit-learn
  • Gradient descent and Ridge
  • Stochastic Gradient Descent
  • Computation of gradients
  • diff --git a/doc/pub/Splines/html/Splines-bs.html b/doc/pub/Splines/html/Splines-bs.html index 09897dcec..e3f254239 100644 --- a/doc/pub/Splines/html/Splines-bs.html +++ b/doc/pub/Splines/html/Splines-bs.html @@ -49,21 +49,21 @@ Automatically generated HTML file from DocOnce source ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -128,18 +128,18 @@ MathJax.Hub.Config({
  • More on Steepest descent
  • The ideal
  • The sensitiveness of the gradient descent
  • -
  • Gradient Descent Example
  • -
  • And a corresponding example using scikit-learn
  • -
  • Convex functions
  • -
  • Convex function
  • -
  • Conditions on convex functions
  • -
  • More on convex functions
  • -
  • Some simple problems
  • -
  • Revisiting our first homework
  • -
  • Gradient descent example
  • -
  • The derivative of the cost/loss function
  • -
  • The Hessian matrix
  • -
  • Simple program
  • +
  • Convex functions
  • +
  • Convex function
  • +
  • Conditions on convex functions
  • +
  • More on convex functions
  • +
  • Some simple problems
  • +
  • Revisiting our first homework
  • +
  • Gradient descent example
  • +
  • The derivative of the cost/loss function
  • +
  • The Hessian matrix
  • +
  • Simple program
  • +
  • Gradient Descent Example
  • +
  • And a corresponding example using scikit-learn
  • Gradient descent and Ridge
  • Stochastic Gradient Descent
  • Computation of gradients
  • @@ -193,7 +193,7 @@ MathJax.Hub.Config({
    [2] Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University

    -

    Sep 20, 2018

    +

    Sep 21, 2018


    diff --git a/doc/pub/Splines/html/Splines-reveal.html b/doc/pub/Splines/html/Splines-reveal.html index f78ef0d25..53afa7c72 100644 --- a/doc/pub/Splines/html/Splines-reveal.html +++ b/doc/pub/Splines/html/Splines-reveal.html @@ -148,7 +148,7 @@ MathJax.Hub.Config({

    [2] Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University

     
    -

    Sep 20, 2018

    +

    Sep 21, 2018


    @@ -186,12 +186,14 @@ direction of the negative gradient \( -\nabla F(\mathbf{x}) \). It can be shown that if

     
    $$ -\mathbf{x}_{k+1} = \mathbf{x}_k - \gamma_k \nabla F(\mathbf{x}_k), \ \ \gamma_k > 0 +\mathbf{x}_{k+1} = \mathbf{x}_k - \gamma_k \nabla F(\mathbf{x}_k), $$

     
    +with \( \gamma_k > 0 \). +

    -for \( \gamma_k \) small enough, then \( F(\mathbf{x}_{k+1}) \leq +For \( \gamma_k \) small enough, then \( F(\mathbf{x}_{k+1}) \leq F(\mathbf{x}_k) \). This means that for a sufficiently small \( \gamma_k \) we are always moving towards smaller function values, i.e a minimum. @@ -222,7 +224,7 @@ the learning rate within the context of Machine Learning.

    The ideal

    -Ideally the sequence \( \{ \mathbf{x}_k \}_{k=0} \) converges to a global +Ideally the sequence \( \{\mathbf{x}_k \}_{k=0} \) converges to a global minimum of the function \( F \). In general we do not know if we are in a global or local minimum. In the special case when \( F \) is a convex function, all local minima are also global minima, so in this case @@ -248,7 +250,8 @@ Note that the gradient is a function of \( \mathbf{x} =

    The sensitiveness of the gradient descent

    -GD is sensitive to the choice of learning rate \( \gamma_k \). This is due +The gradient descent method +is sensitive to the choice of learning rate \( \gamma_k \). This is due to the fact that we are only guaranteed that \( F(\mathbf{x}_{k+1}) \leq F(\mathbf{x}_k) \) for sufficiently small \( \gamma_k \). The problem is to determine an optimal learning rate. If the learning rate is chosen too @@ -263,10 +266,284 @@ randomness. One such method is that of Stochastic Gradient Descent

    -

    Gradient Descent Example

    +

    Convex functions

    -We revisit now our simple linear regression example with a linear polynomial. +Ideally we want our cost/loss function to be convex(concave). + +

    +First we give the definition of a convex set: A set \( C \) in +\( \mathbb{R}^n \) is said to be convex if, for all \( x \) and \( y \) in \( C \) and +all \( t \in (0,1) \) , the point \( (1 − t)x + ty \) also belongs to +C. Geometrically this means that every point on the line segment +connecting \( x \) and \( y \) is in \( C \) as discussed below. + +

    +The convex subsets of \( \mathbb{R} \) are the intervals of +\( \mathbb{R} \). Examples of convex sets of \( \mathbb{R}^2 \) are the +regular polygons (triangles, rectangles, pentagons, etc...). +

    + + +
    +

    Convex function

    + +

    +Convex function: Let \( X \subset \mathbb{R}^n \) be a convex set. Assume that the function \( f: X \rightarrow \mathbb{R} \) is continuous, then \( f \) is said to be convex if

     
    +$$f(tx_1 + (1-t)x_2) \leq tf(x_1) + (1-t)f(x_2) $$ +

     
    for all \( x_1, x_2 \in X \) and for all \( t \in [0,1] \). If \( \leq \) is replaced with a strict inequaltiy in the definition, we demand \( x_1 \neq x_2 \) and \( t\in(0,1) \) then \( f \) is said to be strictly convex. For a single variable function, convexity means that if you draw a straight line connecting \( f(x_1) \) and \( f(x_2) \), the value of the function on the interval \( [x_1,x_2] \) is always below the line as illustrated below. +

    + + +
    +

    Conditions on convex functions

    + +

    +In the following we state first and second-order conditions which +ensures convexity of a function \( f \). We write \( D_f \) to denote the +domain of \( f \), i.e the subset of \( R^n \) where \( f \) is defined. For more +details and proofs we refer to: S. Boyd and L. Vandenberghe. Convex Optimization. Cambridge University Press. + +

    +

    +First order condition. +

    +Suppose \( f \) is differentiable (i.e \( \nabla f(x) \) is well defined for +all \( x \) in the domain of \( f \)). Then \( f \) is convex if and only if \( D_f \) +is a convex set and

     
    +$$f(y) \geq f(x) + \nabla f(x)^T (y-x) $$ +

     
    holds +for all \( x,y \in D_f \). This condition means that for a convex function +the first order Taylor expansion (right hand side above) at any point +a global under estimator of the function. To convince yourself you can +make a drawing of \( f(x) = x^2+1 \) and draw the tangent line to \( f(x) \) and +note that it is always below the graph. +

    + +

    +

    +Second order condition. +

    +Assume that \( f \) is twice +differentiable, i.e the Hessian matrix exists at each point in +\( D_f \). Then \( f \) is convex if and only if \( D_f \) is a convex set and its +Hessian is positive semi-definite for all \( x\in D_f \). For a +single-variable function this reduces to \( f''(x) \geq 0 \). Geometrically this means that \( f \) has nonnegative curvature +everywhere. +

    + +

    +This condition is particularly useful since it gives us an procedure for determining if the function under consideration is convex, apart from using the definition. +

    + + +
    +

    More on convex functions

    + +

    +The next result is of great importance to us and the reason why we are +going on about convex functions. In machine learning we frequently +have to minimize a loss/cost function in order to find the best +parameters for the model we are considering. + +

    +Ideally we want the +global minimum (for high-dimensional models it is hard to know +if we have local or global minimum). However, if the cost/loss function +is convex the following result provides invaluable information: + +

    +

    +Any minimum is global for convex functions. +

    +Consider the problem of finding \( x \in \mathbb{R}^n \) such that \( f(x) \) +is minimal, where \( f \) is convex and differentiable. Then, any point +\( x^* \) that satisfies \( \nabla f(x^*) = 0 \) is a global minimum. +

    + +

    +This result means that if we know that the cost/loss function is convex and we are able to find a minimum, we are guaranteed that it is a global minimum. +

    + + +
    +

    Some simple problems

    + +
      +

    1. Show that \( f(x)=x^2 \) is convex for \( x \in \mathbb{R} \) using the definition of convexity. Hint: If you re-write the definition, \( f \) is convex if the following holds for all \( x,y \in D_f \) and any \( \lambda \in [0,1] \) $\lambda f(x)+(1-\lambda)f(y)-f(\lambda x + (1-\lambda) y ) \geq 0$.
    2. +

    3. Using the second order condition show that the following functions are convex on the specified domain.
    4. + +
        +

      • \( f(x) = e^x \) is convex for \( x \in \mathbb{R} \).
      • +

      • \( g(x) = -\ln(x) \) is convex for \( x \in (0,\infty) \).
      • +
      +

    5. Let \( f(x) = x^2 \) and \( g(x) = e^x \). Show that \( f(g(x)) \) and \( g(f(x)) \) is convex for \( x \in \mathbb{R} \). Also show that if \( f(x) \) is any convex function than \( h(x) = e^{f(x)} \) is convex.
    6. +

    7. A norm is any function that satisfy the following properties
    8. + +
        +

      • \( f(\alpha x) = |\alpha| f(x) \) for all \( \alpha \in \mathbb{R} \).
      • +

      • \( f(x+y) \leq f(x) + f(y) \)
      • +

      • \( f(x) \leq 0 \) for all \( x \in \mathbb{R}^n \) with equality if and only if \( x = 0 \)
      • +
      +

      +

    +

    + +Using the definition of convexity, try to show that a function satisfying the properties above is convex (the third condition is not needed to show this). +

    + + +
    +

    Revisiting our first homework

    + +

    +We will use linear regression as a case study for the gradient descent +methods. Linear regression is a great test case for the gradient +descent methods discussed in the lectures since it has several +desirable properties such as: + +

      +

    1. An analytical solution (recall homework set 1).
    2. +

    3. The gradient can be computed analytically.
    4. +

    5. The cost function is convex which guarantees that gradient descent converges for small enough learning rates
    6. +
    +

    + +We revisit the example from homework set 1 where we had +

     
    +$$ +y_i = 5x_i^2 + 0.1\xi_i, \ i=1,\cdots,100 +$$ +

     
    + +with \( x_i \in [0,1] \) chosen randomly with a uniform distribution. Additionally \( \xi_i \) represents stochastic noise chosen according to a normal distribution \( \cal {N}(0,1) \). +The linear regression model is given by +

     
    +$$ +h_\beta(x) = \hat{y} = \beta_0 + \beta_1 x, +$$ +

     
    + +such that +

     
    +$$ +\hat{y}_i = \beta_0 + \beta_1 x_i. +$$ +

     
    +

    + + +
    +

    Gradient descent example

    + +

    +Let \( \mathbf{y} = (y_1,\cdots,y_n)^T \), \( \mathbf{\hat{y}} = (\hat{y}_1,\cdots,\hat{y}_n)^T \) and \( \beta = (\beta_0, \beta_1)^T \) + +

    +It is convenient to write \( \mathbf{\hat{y}} = X\beta \) where \( X \in \mathbb{R}^{100 \times 2} \) is the design matrix given by +

     
    +$$ +X \equiv \begin{bmatrix} +1 & x_1 \\ +\vdots & \vdots \\ +1 & x_{100} & \\ +\end{bmatrix}. +$$ +

     
    + +The loss function is given by +

     
    +$$ +C(\beta) = ||X\beta-\mathbf{y}||^2 = ||X\beta||^2 - 2 \mathbf{y}^T X\beta + ||\mathbf{y}||^2 = \sum_{i=1}^{100} (\beta_0 + \beta_1 x_i)^2 - 2 y_i (\beta_0 + \beta_1 x_i) + y_i^2 +$$ +

     
    + +and we want to find \( \beta \) such that \( C(\beta) \) is minimized. +

    + + +
    +

    The derivative of the cost/loss function

    + +

    +Computing \( \partial C(\beta) / \partial \beta_0 \) and \( \partial C(\beta) / \partial \beta_1 \) we can show that the gradient can be written as +

     
    +$$ +\nabla_{\beta} C(\beta) = (\partial C(\beta) / \partial \beta_0, \partial C(\beta) / \partial \beta_1)^T = 2\begin{bmatrix} \sum_{i=1}^{100} \left(\beta_0+\beta_1x_i-y_i\right) \\ +\sum_{i=1}^{100}\left( x_i (\beta_0+\beta_1x_i)-y_ix_i\right) \\ +\end{bmatrix} = 2X^T(X\beta - \mathbf{y}), +$$ +

     
    + +where \( X \) is the design matrix defined above. +

    + + +
    +

    The Hessian matrix

    +The Hessian matrix of \( C(\beta) \) is given by +

     
    +$$ +\hat{H} \equiv \begin{bmatrix} +\frac{\partial^2 C(\beta)}{\partial \beta_0^2} & \frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} \\ +\frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} & \frac{\partial^2 C(\beta)}{\partial \beta_1^2} & \\ +\end{bmatrix} = 2X^T X. +$$ +

     
    + +This result implies that \( C(\beta) \) is a convex function since the matrix \( X^T X \) always is positive semi-definite. +

    + + +
    +

    Simple program

    + +

    +We can now write a program that minimizes \( C(\beta) \) using the gradient descent method with a constant learning rate \( \gamma \) according to +

     
    +$$ +\beta_{k+1} = \beta_k - \gamma \nabla_\beta C(\beta_k), \ k=0,1,\cdots +$$ +

     
    + +

    +We can use the expression we computed for the gradient and let use a +\( \beta_0 \) be chosen randomly and let \( \gamma = 0.001 \). Stop iterating +when \( ||\nabla_\beta C(\beta_k) || \leq \epsilon = 10^{-8} \). + +

    +And finally we can compare our solution for \( \beta \) with the analytic result given by +\( \beta= (X^TX)^{-1} X^T \mathbf{y} \). +

    + + +

    import numpy as np
    +
    +"""
    +The following setup is just a suggestion, feel free to write it the way you like.
    +"""
    +
    +#Setup problem described in the exercise
    +N  = 100 #Nr of datapoints
    +M  = 2 #Nr of features
    +x  = np.random.rand(N) #Uniformly generated x-values in [0,1]
    +y  = 5*x**2 + 0.1*np.random.randn(N)
    +X  = np.c_[np.ones(N),x] #Construct design matrix
    +
    +#Compute beta according to normal equations to compare with GD solution
    +Xt_X_inv = np.linalg.inv(np.dot(X.T,X))
    +Xt_y     = np.dot(X.transpose(),y)
    +beta_NE = np.dot(Xt_X_inv,Xt_y)
    +print(beta_NE)
    +
    +
    + + +
    +

    Gradient Descent Example

    + +

    +Another simple example is here

    @@ -313,7 +590,7 @@ plt.show()

    -

    And a corresponding example using scikit-learn

    +

    And a corresponding example using scikit-learn

    @@ -337,284 +614,6 @@ sgdreg.fit(x,y.ravel())

    -
    -

    Convex functions

    - -

    -Ideally we want our cost/loss function to be convex(concave). - -

    -First we give the definition of a convex set: A set \( C \) in -\( \mathbb{R}^n \) is said to be convex if, for all \( x \) and \( y \) in \( C \) and -all \( t \in (0,1) \) , the point \( (1 − t)x + ty \) also belongs to -C. Geometrically this means that every point on the line segment -connecting \( x \) and \( y \) is in \( C \) as discussed below. - -

    -The convex subsets of \( \mathbb{R} \) are the intervals of -\( \mathbb{R} \). Examples of convex sets of \( \mathbb{R}^2 \) are the -regular polygons (triangles, rectangles, pentagons, etc...). -

    - - -
    -

    Convex function

    - -

    -Convex function: Let \( X \subset \mathbb{R}^n \) be a convex set. Assume that the function \( f: X \rightarrow \mathbb{R} \) is continuous, then \( f \) is said to be convex if

     
    -$$f(tx_1 + (1-t)x_2) \leq tf(x_1) + (1-t)f(x_2) $$ -

     
    for all \( x_1, x_2 \in X \) and for all \( t \in [0,1] \). If \( \leq \) is replaced with a strict inequaltiy in the definition, we demand \( x_1 \neq x_2 \) and \( t\in(0,1) \) then \( f \) is said to be strictly convex. For a single variable function, convexity means that if you draw a straight line connecting \( f(x_1) \) and \( f(x_2) \), the value of the function on the interval \( [x_1,x_2] \) is always below the line as illustrated below. -

    - - -
    -

    Conditions on convex functions

    - -

    -In the following we state first and second-order conditions which -ensures convexity of a function \( f \). We write \( D_f \) to denote the -domain of \( f \), i.e the subset of \( R^n \) where \( f \) is defined. For more -details and proofs we refer to: S. Boyd and L. Vandenberghe. Convex Optimization. Cambridge University Press. - -

    -

    -First order condition. -

    -Suppose \( f \) is differentiable (i.e \( \nabla f(x) \) is well defined for -all \( x \) in the domain of \( f \)). Then \( f \) is convex if and only if \( D_f \) -is a convex set and

     
    -$$f(y) \geq f(x) + \nabla f(x)^T (y-x) $$ -

     
    holds -for all \( x,y \in D_f \). This condition means that for a convex function -the first order Taylor expansion (right hand side above) at any point -a global under estimator of the function. To convince yourself you can -make a drawing of f(x) = x^2+1 and draw the tangent line to \( f(x) \) and -note that it is always below the graph. -

    - -

    -

    -Second order condition. -

    -Assume that \( f \) is twice -differentiable, i.e the Hessian matrix exists at each point in -\( D_f \). Then \( f \) is convex if and only if \( D_f \) is a convex set and its -Hessian is positive semi-definite for all \( x\in D_f \). For a -single-variable function this reduces to \( f''(x) \geq -0 \). Geometrically this means that \( f \) has nonnegative curvature -everywhere. -

    - -

    -This condition is particularly useful since it gives us an procedure for determining if the function under consideration is convex, apart from using the definition. -

    - - -
    -

    More on convex functions

    - -

    -The next result is of great importance to us and the reason why we are -going on about convex functions. In machine learning we frequently -have to minimize a loss/cost function in order to find the best -parameters for the model we are considering. - -

    -Ideally we want the -global minimum (for high-dimensional models it is hard to know -if we have local or global minimum). However, if the cost/loss function -is convex the following result provides invaluable information: - -

    -

    -Any minimum is global for convex functions. -

    -Consider the problem of finding \( x \in \mathbb{R}^n \) such that \( f(x) \) -is minimal, where \( f \) is convex and differentiable. Then, any point -\( x^* \) that satisfies \( \nabla f(x^*) = 0 \) is a global minimum. -

    - -

    -This result means that if we know that the cost/loss function is convex and we are able to find a minimum, we are guaranteed that it is a global minimum. -

    - - -
    -

    Some simple problems

    - -
      -

    1. Show that \( f(x)=x^2 \) is convex for \( x \in \mathbb{R} \) using the definition of convexity. Hint: If you re-write the definition, \( f \) is convex if the following holds for all \( x,y \in D_f \) and any \( \lambda \in [0,1] \) $\lambda f(x) + (1-\lambda)f(y) - f(\lambda x + (1-\lambda) y ) \geq 0. $
    2. -

    3. Using the second order condition show that the following functions are convex on the specified domain.
    4. - -
        -

      • \( f(x) = e^x \) is convex for \( x \in \mathbb{R} \).
      • -

      • \( g(x) = -\ln(x) \) is convex for \( x \in (0,\infty) \).
      • -
      -

    5. Let \( f(x) = x^2 \) and \( g(x) = e^x \). Show that \( f(g(x)) \) and \( g(f(x)) \) is convex for \( x \in \mathbb{R} \). Also show that if \( f(x) \) is any convex function than \( h(x) = e^{f(x)} \) is convex.
    6. -

    7. A norm is any function that satisfy the following properties
    8. - -
        -

      • \( f(\alpha x) = |\alpha| f(x) \) for all \( \alpha \in \mathbb{R} \).
      • -

      • \( f(x+y) \leq f(x) + f(y) \)
      • -

      • \( f(x) \leq 0 \) for all \( x \in \mathbb{R}^n \) with equality if and only if \( x = 0 \)
      • -
      -

      -

    -

    - -Using the definition of convexity, try to show that a function satisfying the properties above is convex (the third condition is not needed to show this). -

    - - -
    -

    Revisiting our first homework

    - -

    -We will use linear regression as a case study for the gradient descent -methods. Linear regression is a great test case for the gradient -descent methods discussed in the lectures since it has several -desirable properties such as: - -

      -

    1. An analytical solution (recall homework set 1).
    2. -

    3. The gradient can be computed analytically.
    4. -

    5. The cost function is convex which guarantees that gradient descent converges for small enough learning rates
    6. -
    -

    - -We revisit the example from homework set 1 where we had -

     
    -$$ -y_i = 5x_i^2 + 0.1\xi_i, \ i=1,\cdots,100 -$$ -

     
    - -with \( x_i \in [0,1] \) chosen randomly with a uniform distribution. Additionally \( \xi_i \) represents stochastic noise chosen according to a normal distribution \( \cal {N}(0,1) \). -The linear regression model is given by -

     
    -$$ -h_\beta(x) = \hat{y} = \beta_0 + \beta_1 x, -$$ -

     
    - -such that -

     
    -$$ -\hat{y}_i = \beta_0 + \beta_1 x_i. -$$ -

     
    -

    - - -
    -

    Gradient descent example

    - -

    -Let \( \mathbf{y} = (y_1,\cdots,y_n)^T \), \( \mathbf{\hat{y}} = (\hat{y}_1,\cdots,\hat{y}_n)^T \) and \( \beta = (\beta_0, \beta_1)^T \) - -

    -t is convenient to write \( \mathbf{\hat{y}} = X\beta \) where \( X \in \mathbb{R}^{100 \times 2} \) is the design matrix given by -

     
    -$$ -\begin{equation} -X \equiv \begin{bmatrix} -1 & x_1 \\ -\vdots & \vdots \\ -1 & x_{100} & \\ -\end{bmatrix}. -\tag{1} -\end{equation} -$$ -

     
    - -The loss function is given by -

     
    -$$ -C(\beta) = ||X\beta-\mathbf{y}||^2 = ||X\beta||^2 - 2 \mathbf{y}^T X\beta + ||\mathbf{y}||^2 = \sum_{i=1}^{100} (\beta_0 + \beta_1 x_i)^2 - 2 y_i (\beta_0 + \beta_1 x_i) + y_i^2 -$$ -

     
    - -and we want to find \( \beta \) such that \( C(\beta) \) is minimized. -

    - - -
    -

    The derivative of the cost/loss function

    - -

    -Computing \( \partial C(\beta) / \partial \beta_0 \) and \( \partial C(\beta) / \partial \beta_1 \) we can show that the gradient can be written as -

     
    -$$ -\nabla_\beta C(\beta) = (\partial C(\beta) / \partial \beta_0, \partial C(\beta) / \partial \beta_1)^T = 2\begin{bmatrix} \sum_{i=1}^{100} \left(\beta_0+\beta_1x_i-y_i\right) \\ -\sum_{i=1}^{100}\left( x_i (\beta_0+\beta_1x_i)-y_ix_i\right) \\ -\end{bmatrix} = 2X^T(X\beta - \mathbf{y}), -$$ -

     
    - -where \( X \) is the design matrix defined above. -

    - - -
    -

    The Hessian matrix

    -The Hessian matrix of \( C(\beta) \) is given by -

     
    -$$ -\hat{H} \equiv \begin{bmatrix} -\frac{\partial^2 C(\beta)}{\partial \beta_0^2} & \frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} \\ -\frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} & \frac{\partial^2 C(\beta)}{\partial \beta_1^2} & \\ -\end{bmatrix} = 2X^T X. -$$ -

     
    - -This result implies that \( C(\beta) \) is a convex function since the matrix \( X^T X \) always is positive semi-definite. -

    - - -
    -

    Simple program

    - -

    -We can now write a program that minimizes \( C(\beta) \) using the gradient descent method with a constant learning rate \( \gamma \) according to -

     
    -$$ -\beta_{k+1} = \beta_k - \gamma \nabla_\beta C(\beta_k), \ k=0,1,\cdots -$$ -

     
    - -

    -We can use the expression we computed for the gradient and let use a -\( \beta_0 \) be chosen randomly and let \( \gamma = 0.001 \). Stop iterating -when \( ||\nabla_\beta C(\beta_k) || < \epsilon = 10^{-8} \). - -

    -And finally we can compare our solution for \( \beta \) with the analytic result given by -\( \beta= (X^TX)^{-1} X^T \mathbf{y} \). -

    - - -

    import numpy as np
    -
    -"""
    -The following setup is just a suggestion, feel free to write it the way you like.
    -"""
    -
    -#Setup problem described in the exercise
    -N  = 100 #Nr of datapoints
    -M  = 2 #Nr of features
    -x  = np.random.rand(N) #Uniformly generated x-values in [0,1]
    -y  = 5*x**2 + 0.1*np.random.randn(N)
    -X  = np.c_[np.ones(N),x] #Construct design matrix
    -
    -#Compute beta according to normal equations to compare with GD solution
    -Xt_X_inv = np.linalg.inv(np.dot(X.T,X))
    -Xt_y     = np.dot(X.transpose(),y)
    -beta_NE = np.dot(Xt_X_inv,Xt_y)
    -print(beta_NE)
    -
    -
    - -

    Gradient descent and Ridge

    @@ -689,7 +688,7 @@ the shortcomings of the Gradient descent method discussed above.

    The underlying idea of SGD comes from the observation that the cost function, which we want to minimize, can almost always be written as a -sum over \( n \) datapoints \( \{\mathbf{x}_i\}_{i=1}^n \), +sum over \( n \) data points \( \{\mathbf{x}_i\}_{i=1}^n \),

     
    $$ C(\mathbf{\beta}) = \sum_{i=1}^n c_i(\mathbf{x}_i, @@ -715,7 +714,7 @@ $$

    Stochasticity/randomness is introduced by only taking the gradient on a subset of the data called minibatches. If there are \( n \) -datapoints and the size of each minibatch is \( M \), there will be \( n/M \) +data points and the size of each minibatch is \( M \), there will be \( n/M \) minibatches. We denote these minibatches by \( B_k \) where \( k=1,\cdots,n/M \).

    @@ -723,22 +722,22 @@ minibatches. We denote these minibatches by \( B_k \) where

    SGD example

    -As an example, suppose we have \( 10 \) datapoints \( ( \mathbf{x}_1, -\cdots, \mathbf{x}_{10} ) \) and we choose to have \( M=5 \) minibathces, -then each minibatch contains two datapoints. In particular we have +As an example, suppose we have \( 10 \) data points \( (\mathbf{x}_1,\cdots, \mathbf{x}_{10}) \) +and we choose to have \( M=5 \) minibathces, +then each minibatch contains two data points. In particular we have \( B_1 = (\mathbf{x}_1,\mathbf{x}_2), \cdots, B_5 = (\mathbf{x}_9,\mathbf{x}_{10}) \). Note that if you choose \( M=1 \) you -have only a single batch with all datapoints and on the other extreme, +have only a single batch with all data points and on the other extreme, you may choose \( M=n \) resulting in a minibatch for each datapoint, i.e \( B_k = \mathbf{x}_k \).

    The idea is now to approximate the gradient by replacing the sum over -all datapoints with a sum over the datapoints in one the minibatches +all data points with a sum over the data points in one the minibatches picked at random in each gradient descent step

     
    $$ -\nabla_\beta +\nabla_{\beta} C(\mathbf{\beta}) = \sum_{i=1}^n \nabla_\beta c_i(\mathbf{x}_i, \mathbf{\beta}) \rightarrow \sum_{i \in B_k}^n \nabla_\beta c_i(\mathbf{x}_i, \mathbf{\beta}). @@ -794,8 +793,8 @@ Taking the gradient only on a subset of the data has two important benefits. First, it introduces randomness which decreases the chance that our opmization scheme gets stuck in a local minima. Second, if the size of the minibatches are small relative to the number of -datapoints (\( M < n \)), the computation of the gradient is much -cheaper since we sum over the datapoints in the k-th minibatch and not +datapoints (\( M < n \)), the computation of the gradient is much +cheaper since we sum over the datapoints in the \( k-th \) minibatch and not all \( n \) datapoints.

    @@ -826,7 +825,7 @@ number of epochs in such a way that it becomes very small after a reasonable time such that we do not move at all.

    -As an example, let \( e = 0,1,2,3,\cdots \) denote the current epoch and let \( t_0, t_1 > 0 \) be two fixed numbers. Furthermore, let \( t = e \cdot m + i \) where \( m \) is the number of minibatches and \( i=0,\cdots,m-1 \). Then the function

     
    +As an example, let \( e = 0,1,2,3,\cdots \) denote the current epoch and let \( t_0, t_1 > 0 \) be two fixed numbers. Furthermore, let \( t = e \cdot m + i \) where \( m \) is the number of minibatches and \( i=0,\cdots,m-1 \). Then the function

     
    $$\gamma_j(t; t_0, t_1) = \frac{t_0}{t+t_1} $$

     
    goes to zero as the number of epochs gets large. I.e. we start with a step length \( \gamma_j (0; t_0, t_1) = t_0/t_1 \) which decays in time \( t \). diff --git a/doc/pub/Splines/html/Splines-solarized.html b/doc/pub/Splines/html/Splines-solarized.html index fc0bfdc82..752bc03f2 100644 --- a/doc/pub/Splines/html/Splines-solarized.html +++ b/doc/pub/Splines/html/Splines-solarized.html @@ -69,21 +69,21 @@ div { text-align: justify; text-justify: inter-word; } ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -147,7 +147,7 @@ MathJax.Hub.Config({

    [2] Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University

    -

    Sep 20, 2018

    +

    Sep 21, 2018












    @@ -178,11 +178,13 @@ direction of the negative gradient \( -\nabla F(\mathbf{x}) \).

    It can be shown that if $$ -\mathbf{x}_{k+1} = \mathbf{x}_k - \gamma_k \nabla F(\mathbf{x}_k), \ \ \gamma_k > 0 +\mathbf{x}_{k+1} = \mathbf{x}_k - \gamma_k \nabla F(\mathbf{x}_k), $$ +with \( \gamma_k > 0 \). +

    -for \( \gamma_k \) small enough, then \( F(\mathbf{x}_{k+1}) \leq +For \( \gamma_k \) small enough, then \( F(\mathbf{x}_{k+1}) \leq F(\mathbf{x}_k) \). This means that for a sufficiently small \( \gamma_k \) we are always moving towards smaller function values, i.e a minimum. @@ -211,7 +213,7 @@ the learning rate within the context of Machine Learning.

    The ideal

    -Ideally the sequence \( \{ \mathbf{x}_k \}_{k=0} \) converges to a global +Ideally the sequence \( \{\mathbf{x}_k \}_{k=0} \) converges to a global minimum of the function \( F \). In general we do not know if we are in a global or local minimum. In the special case when \( F \) is a convex function, all local minima are also global minima, so in this case @@ -237,7 +239,8 @@ Note that the gradient is a function of \( \mathbf{x} =

    The sensitiveness of the gradient descent

    -GD is sensitive to the choice of learning rate \( \gamma_k \). This is due +The gradient descent method +is sensitive to the choice of learning rate \( \gamma_k \). This is due to the fact that we are only guaranteed that \( F(\mathbf{x}_{k+1}) \leq F(\mathbf{x}_k) \) for sufficiently small \( \gamma_k \). The problem is to determine an optimal learning rate. If the learning rate is chosen too @@ -250,12 +253,267 @@ randomness. One such method is that of Stochastic Gradient Descent (SGD), see below.

    -









    + -

    Gradient Descent Example

    +

    Convex functions

    -We revisit now our simple linear regression example with a linear polynomial. +Ideally we want our cost/loss function to be convex(concave). + +

    +First we give the definition of a convex set: A set \( C \) in +\( \mathbb{R}^n \) is said to be convex if, for all \( x \) and \( y \) in \( C \) and +all \( t \in (0,1) \) , the point \( (1 − t)x + ty \) also belongs to +C. Geometrically this means that every point on the line segment +connecting \( x \) and \( y \) is in \( C \) as discussed below. + +

    +The convex subsets of \( \mathbb{R} \) are the intervals of +\( \mathbb{R} \). Examples of convex sets of \( \mathbb{R}^2 \) are the +regular polygons (triangles, rectangles, pentagons, etc...). + +

    +









    + +

    Convex function

    + +

    +Convex function: Let \( X \subset \mathbb{R}^n \) be a convex set. Assume that the function \( f: X \rightarrow \mathbb{R} \) is continuous, then \( f \) is said to be convex if $$f(tx_1 + (1-t)x_2) \leq tf(x_1) + (1-t)f(x_2) $$ for all \( x_1, x_2 \in X \) and for all \( t \in [0,1] \). If \( \leq \) is replaced with a strict inequaltiy in the definition, we demand \( x_1 \neq x_2 \) and \( t\in(0,1) \) then \( f \) is said to be strictly convex. For a single variable function, convexity means that if you draw a straight line connecting \( f(x_1) \) and \( f(x_2) \), the value of the function on the interval \( [x_1,x_2] \) is always below the line as illustrated below. + +

    +









    + +

    Conditions on convex functions

    + +

    +In the following we state first and second-order conditions which +ensures convexity of a function \( f \). We write \( D_f \) to denote the +domain of \( f \), i.e the subset of \( R^n \) where \( f \) is defined. For more +details and proofs we refer to: S. Boyd and L. Vandenberghe. Convex Optimization. Cambridge University Press. + +

    +

    +First order condition. +

    +Suppose \( f \) is differentiable (i.e \( \nabla f(x) \) is well defined for +all \( x \) in the domain of \( f \)). Then \( f \) is convex if and only if \( D_f \) +is a convex set and $$f(y) \geq f(x) + \nabla f(x)^T (y-x) $$ holds +for all \( x,y \in D_f \). This condition means that for a convex function +the first order Taylor expansion (right hand side above) at any point +a global under estimator of the function. To convince yourself you can +make a drawing of \( f(x) = x^2+1 \) and draw the tangent line to \( f(x) \) and +note that it is always below the graph. +

    + + +

    +

    +Second order condition. +

    +Assume that \( f \) is twice +differentiable, i.e the Hessian matrix exists at each point in +\( D_f \). Then \( f \) is convex if and only if \( D_f \) is a convex set and its +Hessian is positive semi-definite for all \( x\in D_f \). For a +single-variable function this reduces to \( f''(x) \geq 0 \). Geometrically this means that \( f \) has nonnegative curvature +everywhere. +

    + + +

    +This condition is particularly useful since it gives us an procedure for determining if the function under consideration is convex, apart from using the definition. + +

    +









    + +

    More on convex functions

    + +

    +The next result is of great importance to us and the reason why we are +going on about convex functions. In machine learning we frequently +have to minimize a loss/cost function in order to find the best +parameters for the model we are considering. + +

    +Ideally we want the +global minimum (for high-dimensional models it is hard to know +if we have local or global minimum). However, if the cost/loss function +is convex the following result provides invaluable information: + +

    +

    +Any minimum is global for convex functions. +

    +Consider the problem of finding \( x \in \mathbb{R}^n \) such that \( f(x) \) +is minimal, where \( f \) is convex and differentiable. Then, any point +\( x^* \) that satisfies \( \nabla f(x^*) = 0 \) is a global minimum. +

    + + +

    +This result means that if we know that the cost/loss function is convex and we are able to find a minimum, we are guaranteed that it is a global minimum. + +

    +









    + +

    Some simple problems

    + +
      +
    1. Show that \( f(x)=x^2 \) is convex for \( x \in \mathbb{R} \) using the definition of convexity. Hint: If you re-write the definition, \( f \) is convex if the following holds for all \( x,y \in D_f \) and any \( \lambda \in [0,1] \) $\lambda f(x)+(1-\lambda)f(y)-f(\lambda x + (1-\lambda) y ) \geq 0$.
    2. +
    3. Using the second order condition show that the following functions are convex on the specified domain.
    4. + +
        +
      • \( f(x) = e^x \) is convex for \( x \in \mathbb{R} \).
      • +
      • \( g(x) = -\ln(x) \) is convex for \( x \in (0,\infty) \).
      • +
      + +
    5. Let \( f(x) = x^2 \) and \( g(x) = e^x \). Show that \( f(g(x)) \) and \( g(f(x)) \) is convex for \( x \in \mathbb{R} \). Also show that if \( f(x) \) is any convex function than \( h(x) = e^{f(x)} \) is convex.
    6. +
    7. A norm is any function that satisfy the following properties
    8. + +
        +
      • \( f(\alpha x) = |\alpha| f(x) \) for all \( \alpha \in \mathbb{R} \).
      • +
      • \( f(x+y) \leq f(x) + f(y) \)
      • +
      • \( f(x) \leq 0 \) for all \( x \in \mathbb{R}^n \) with equality if and only if \( x = 0 \)
      • +
      + +
    + +Using the definition of convexity, try to show that a function satisfying the properties above is convex (the third condition is not needed to show this). + +

    + + +

    Revisiting our first homework

    + +

    +We will use linear regression as a case study for the gradient descent +methods. Linear regression is a great test case for the gradient +descent methods discussed in the lectures since it has several +desirable properties such as: + +

      +
    1. An analytical solution (recall homework set 1).
    2. +
    3. The gradient can be computed analytically.
    4. +
    5. The cost function is convex which guarantees that gradient descent converges for small enough learning rates
    6. +
    + +We revisit the example from homework set 1 where we had +$$ +y_i = 5x_i^2 + 0.1\xi_i, \ i=1,\cdots,100 +$$ + +with \( x_i \in [0,1] \) chosen randomly with a uniform distribution. Additionally \( \xi_i \) represents stochastic noise chosen according to a normal distribution \( \cal {N}(0,1) \). +The linear regression model is given by +$$ +h_\beta(x) = \hat{y} = \beta_0 + \beta_1 x, +$$ + +such that +$$ +\hat{y}_i = \beta_0 + \beta_1 x_i. +$$ + +

    + + +

    Gradient descent example

    + +

    +Let \( \mathbf{y} = (y_1,\cdots,y_n)^T \), \( \mathbf{\hat{y}} = (\hat{y}_1,\cdots,\hat{y}_n)^T \) and \( \beta = (\beta_0, \beta_1)^T \) + +

    +It is convenient to write \( \mathbf{\hat{y}} = X\beta \) where \( X \in \mathbb{R}^{100 \times 2} \) is the design matrix given by +$$ +X \equiv \begin{bmatrix} +1 & x_1 \\ +\vdots & \vdots \\ +1 & x_{100} & \\ +\end{bmatrix}. +$$ + +The loss function is given by +$$ +C(\beta) = ||X\beta-\mathbf{y}||^2 = ||X\beta||^2 - 2 \mathbf{y}^T X\beta + ||\mathbf{y}||^2 = \sum_{i=1}^{100} (\beta_0 + \beta_1 x_i)^2 - 2 y_i (\beta_0 + \beta_1 x_i) + y_i^2 +$$ + +and we want to find \( \beta \) such that \( C(\beta) \) is minimized. + +

    +









    + +

    The derivative of the cost/loss function

    + +

    +Computing \( \partial C(\beta) / \partial \beta_0 \) and \( \partial C(\beta) / \partial \beta_1 \) we can show that the gradient can be written as +$$ +\nabla_{\beta} C(\beta) = (\partial C(\beta) / \partial \beta_0, \partial C(\beta) / \partial \beta_1)^T = 2\begin{bmatrix} \sum_{i=1}^{100} \left(\beta_0+\beta_1x_i-y_i\right) \\ +\sum_{i=1}^{100}\left( x_i (\beta_0+\beta_1x_i)-y_ix_i\right) \\ +\end{bmatrix} = 2X^T(X\beta - \mathbf{y}), +$$ + +where \( X \) is the design matrix defined above. + +

    +









    + +

    The Hessian matrix

    +The Hessian matrix of \( C(\beta) \) is given by +$$ +\hat{H} \equiv \begin{bmatrix} +\frac{\partial^2 C(\beta)}{\partial \beta_0^2} & \frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} \\ +\frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} & \frac{\partial^2 C(\beta)}{\partial \beta_1^2} & \\ +\end{bmatrix} = 2X^T X. +$$ + +This result implies that \( C(\beta) \) is a convex function since the matrix \( X^T X \) always is positive semi-definite. + +

    +









    + +

    Simple program

    + +

    +We can now write a program that minimizes \( C(\beta) \) using the gradient descent method with a constant learning rate \( \gamma \) according to +$$ +\beta_{k+1} = \beta_k - \gamma \nabla_\beta C(\beta_k), \ k=0,1,\cdots +$$ + +

    +We can use the expression we computed for the gradient and let use a +\( \beta_0 \) be chosen randomly and let \( \gamma = 0.001 \). Stop iterating +when \( ||\nabla_\beta C(\beta_k) || \leq \epsilon = 10^{-8} \). + +

    +And finally we can compare our solution for \( \beta \) with the analytic result given by +\( \beta= (X^TX)^{-1} X^T \mathbf{y} \). +

    + + +

    import numpy as np
    +
    +"""
    +The following setup is just a suggestion, feel free to write it the way you like.
    +"""
    +
    +#Setup problem described in the exercise
    +N  = 100 #Nr of datapoints
    +M  = 2 #Nr of features
    +x  = np.random.rand(N) #Uniformly generated x-values in [0,1]
    +y  = 5*x**2 + 0.1*np.random.randn(N)
    +X  = np.c_[np.ones(N),x] #Construct design matrix
    +
    +#Compute beta according to normal equations to compare with GD solution
    +Xt_X_inv = np.linalg.inv(np.dot(X.T,X))
    +Xt_y     = np.dot(X.transpose(),y)
    +beta_NE = np.dot(Xt_X_inv,Xt_y)
    +print(beta_NE)
    +
    +

    +









    + +

    Gradient Descent Example

    + +

    +Another simple example is here

    @@ -301,7 +559,7 @@ plt.show()











    -

    And a corresponding example using scikit-learn

    +

    And a corresponding example using scikit-learn

    @@ -325,265 +583,6 @@ sgdreg.fit(x,y.ravel())

    -

    Convex functions

    - -

    -Ideally we want our cost/loss function to be convex(concave). - -

    -First we give the definition of a convex set: A set \( C \) in -\( \mathbb{R}^n \) is said to be convex if, for all \( x \) and \( y \) in \( C \) and -all \( t \in (0,1) \) , the point \( (1 − t)x + ty \) also belongs to -C. Geometrically this means that every point on the line segment -connecting \( x \) and \( y \) is in \( C \) as discussed below. - -

    -The convex subsets of \( \mathbb{R} \) are the intervals of -\( \mathbb{R} \). Examples of convex sets of \( \mathbb{R}^2 \) are the -regular polygons (triangles, rectangles, pentagons, etc...). - -

    -









    - -

    Convex function

    - -

    -Convex function: Let \( X \subset \mathbb{R}^n \) be a convex set. Assume that the function \( f: X \rightarrow \mathbb{R} \) is continuous, then \( f \) is said to be convex if $$f(tx_1 + (1-t)x_2) \leq tf(x_1) + (1-t)f(x_2) $$ for all \( x_1, x_2 \in X \) and for all \( t \in [0,1] \). If \( \leq \) is replaced with a strict inequaltiy in the definition, we demand \( x_1 \neq x_2 \) and \( t\in(0,1) \) then \( f \) is said to be strictly convex. For a single variable function, convexity means that if you draw a straight line connecting \( f(x_1) \) and \( f(x_2) \), the value of the function on the interval \( [x_1,x_2] \) is always below the line as illustrated below. - -

    -









    - -

    Conditions on convex functions

    - -

    -In the following we state first and second-order conditions which -ensures convexity of a function \( f \). We write \( D_f \) to denote the -domain of \( f \), i.e the subset of \( R^n \) where \( f \) is defined. For more -details and proofs we refer to: S. Boyd and L. Vandenberghe. Convex Optimization. Cambridge University Press. - -

    -

    -First order condition. -

    -Suppose \( f \) is differentiable (i.e \( \nabla f(x) \) is well defined for -all \( x \) in the domain of \( f \)). Then \( f \) is convex if and only if \( D_f \) -is a convex set and $$f(y) \geq f(x) + \nabla f(x)^T (y-x) $$ holds -for all \( x,y \in D_f \). This condition means that for a convex function -the first order Taylor expansion (right hand side above) at any point -a global under estimator of the function. To convince yourself you can -make a drawing of f(x) = x^2+1 and draw the tangent line to \( f(x) \) and -note that it is always below the graph. -

    - - -

    -

    -Second order condition. -

    -Assume that \( f \) is twice -differentiable, i.e the Hessian matrix exists at each point in -\( D_f \). Then \( f \) is convex if and only if \( D_f \) is a convex set and its -Hessian is positive semi-definite for all \( x\in D_f \). For a -single-variable function this reduces to \( f''(x) \geq -0 \). Geometrically this means that \( f \) has nonnegative curvature -everywhere. -

    - - -

    -This condition is particularly useful since it gives us an procedure for determining if the function under consideration is convex, apart from using the definition. - -

    -









    - -

    More on convex functions

    - -

    -The next result is of great importance to us and the reason why we are -going on about convex functions. In machine learning we frequently -have to minimize a loss/cost function in order to find the best -parameters for the model we are considering. - -

    -Ideally we want the -global minimum (for high-dimensional models it is hard to know -if we have local or global minimum). However, if the cost/loss function -is convex the following result provides invaluable information: - -

    -

    -Any minimum is global for convex functions. -

    -Consider the problem of finding \( x \in \mathbb{R}^n \) such that \( f(x) \) -is minimal, where \( f \) is convex and differentiable. Then, any point -\( x^* \) that satisfies \( \nabla f(x^*) = 0 \) is a global minimum. -

    - - -

    -This result means that if we know that the cost/loss function is convex and we are able to find a minimum, we are guaranteed that it is a global minimum. - -

    -









    - -

    Some simple problems

    - -
      -
    1. Show that \( f(x)=x^2 \) is convex for \( x \in \mathbb{R} \) using the definition of convexity. Hint: If you re-write the definition, \( f \) is convex if the following holds for all \( x,y \in D_f \) and any \( \lambda \in [0,1] \) $\lambda f(x) + (1-\lambda)f(y) - f(\lambda x + (1-\lambda) y ) \geq 0. $
    2. -
    3. Using the second order condition show that the following functions are convex on the specified domain.
    4. - -
        -
      • \( f(x) = e^x \) is convex for \( x \in \mathbb{R} \).
      • -
      • \( g(x) = -\ln(x) \) is convex for \( x \in (0,\infty) \).
      • -
      - -
    5. Let \( f(x) = x^2 \) and \( g(x) = e^x \). Show that \( f(g(x)) \) and \( g(f(x)) \) is convex for \( x \in \mathbb{R} \). Also show that if \( f(x) \) is any convex function than \( h(x) = e^{f(x)} \) is convex.
    6. -
    7. A norm is any function that satisfy the following properties
    8. - -
        -
      • \( f(\alpha x) = |\alpha| f(x) \) for all \( \alpha \in \mathbb{R} \).
      • -
      • \( f(x+y) \leq f(x) + f(y) \)
      • -
      • \( f(x) \leq 0 \) for all \( x \in \mathbb{R}^n \) with equality if and only if \( x = 0 \)
      • -
      - -
    - -Using the definition of convexity, try to show that a function satisfying the properties above is convex (the third condition is not needed to show this). - -

    - - -

    Revisiting our first homework

    - -

    -We will use linear regression as a case study for the gradient descent -methods. Linear regression is a great test case for the gradient -descent methods discussed in the lectures since it has several -desirable properties such as: - -

      -
    1. An analytical solution (recall homework set 1).
    2. -
    3. The gradient can be computed analytically.
    4. -
    5. The cost function is convex which guarantees that gradient descent converges for small enough learning rates
    6. -
    - -We revisit the example from homework set 1 where we had -$$ -y_i = 5x_i^2 + 0.1\xi_i, \ i=1,\cdots,100 -$$ - -with \( x_i \in [0,1] \) chosen randomly with a uniform distribution. Additionally \( \xi_i \) represents stochastic noise chosen according to a normal distribution \( \cal {N}(0,1) \). -The linear regression model is given by -$$ -h_\beta(x) = \hat{y} = \beta_0 + \beta_1 x, -$$ - -such that -$$ -\hat{y}_i = \beta_0 + \beta_1 x_i. -$$ - -

    - - -

    Gradient descent example

    - -

    -Let \( \mathbf{y} = (y_1,\cdots,y_n)^T \), \( \mathbf{\hat{y}} = (\hat{y}_1,\cdots,\hat{y}_n)^T \) and \( \beta = (\beta_0, \beta_1)^T \) - -

    -t is convenient to write \( \mathbf{\hat{y}} = X\beta \) where \( X \in \mathbb{R}^{100 \times 2} \) is the design matrix given by -$$ -\begin{equation} -X \equiv \begin{bmatrix} -1 & x_1 \\ -\vdots & \vdots \\ -1 & x_{100} & \\ -\end{bmatrix}. -\label{_auto1} -\end{equation} -$$ - -The loss function is given by -$$ -C(\beta) = ||X\beta-\mathbf{y}||^2 = ||X\beta||^2 - 2 \mathbf{y}^T X\beta + ||\mathbf{y}||^2 = \sum_{i=1}^{100} (\beta_0 + \beta_1 x_i)^2 - 2 y_i (\beta_0 + \beta_1 x_i) + y_i^2 -$$ - -and we want to find \( \beta \) such that \( C(\beta) \) is minimized. - -

    -









    - -

    The derivative of the cost/loss function

    - -

    -Computing \( \partial C(\beta) / \partial \beta_0 \) and \( \partial C(\beta) / \partial \beta_1 \) we can show that the gradient can be written as -$$ -\nabla_\beta C(\beta) = (\partial C(\beta) / \partial \beta_0, \partial C(\beta) / \partial \beta_1)^T = 2\begin{bmatrix} \sum_{i=1}^{100} \left(\beta_0+\beta_1x_i-y_i\right) \\ -\sum_{i=1}^{100}\left( x_i (\beta_0+\beta_1x_i)-y_ix_i\right) \\ -\end{bmatrix} = 2X^T(X\beta - \mathbf{y}), -$$ - -where \( X \) is the design matrix defined above. - -

    -









    - -

    The Hessian matrix

    -The Hessian matrix of \( C(\beta) \) is given by -$$ -\hat{H} \equiv \begin{bmatrix} -\frac{\partial^2 C(\beta)}{\partial \beta_0^2} & \frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} \\ -\frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} & \frac{\partial^2 C(\beta)}{\partial \beta_1^2} & \\ -\end{bmatrix} = 2X^T X. -$$ - -This result implies that \( C(\beta) \) is a convex function since the matrix \( X^T X \) always is positive semi-definite. - -

    -









    - -

    Simple program

    - -

    -We can now write a program that minimizes \( C(\beta) \) using the gradient descent method with a constant learning rate \( \gamma \) according to -$$ -\beta_{k+1} = \beta_k - \gamma \nabla_\beta C(\beta_k), \ k=0,1,\cdots -$$ - -

    -We can use the expression we computed for the gradient and let use a -\( \beta_0 \) be chosen randomly and let \( \gamma = 0.001 \). Stop iterating -when \( ||\nabla_\beta C(\beta_k) || < \epsilon = 10^{-8} \). - -

    -And finally we can compare our solution for \( \beta \) with the analytic result given by -\( \beta= (X^TX)^{-1} X^T \mathbf{y} \). -

    - - -

    import numpy as np
    -
    -"""
    -The following setup is just a suggestion, feel free to write it the way you like.
    -"""
    -
    -#Setup problem described in the exercise
    -N  = 100 #Nr of datapoints
    -M  = 2 #Nr of features
    -x  = np.random.rand(N) #Uniformly generated x-values in [0,1]
    -y  = 5*x**2 + 0.1*np.random.randn(N)
    -X  = np.c_[np.ones(N),x] #Construct design matrix
    -
    -#Compute beta according to normal equations to compare with GD solution
    -Xt_X_inv = np.linalg.inv(np.dot(X.T,X))
    -Xt_y     = np.dot(X.transpose(),y)
    -beta_NE = np.dot(Xt_X_inv,Xt_y)
    -print(beta_NE)
    -
    -

    - -

    Gradient descent and Ridge

    @@ -650,7 +649,7 @@ the shortcomings of the Gradient descent method discussed above.

    The underlying idea of SGD comes from the observation that the cost function, which we want to minimize, can almost always be written as a -sum over \( n \) datapoints \( \{\mathbf{x}_i\}_{i=1}^n \), +sum over \( n \) data points \( \{\mathbf{x}_i\}_{i=1}^n \), $$ C(\mathbf{\beta}) = \sum_{i=1}^n c_i(\mathbf{x}_i, \mathbf{\beta}). @@ -672,7 +671,7 @@ $$

    Stochasticity/randomness is introduced by only taking the gradient on a subset of the data called minibatches. If there are \( n \) -datapoints and the size of each minibatch is \( M \), there will be \( n/M \) +data points and the size of each minibatch is \( M \), there will be \( n/M \) minibatches. We denote these minibatches by \( B_k \) where \( k=1,\cdots,n/M \). @@ -680,21 +679,21 @@ minibatches. We denote these minibatches by \( B_k \) where









    SGD example

    -As an example, suppose we have \( 10 \) datapoints \( ( \mathbf{x}_1, -\cdots, \mathbf{x}_{10} ) \) and we choose to have \( M=5 \) minibathces, -then each minibatch contains two datapoints. In particular we have +As an example, suppose we have \( 10 \) data points \( (\mathbf{x}_1,\cdots, \mathbf{x}_{10}) \) +and we choose to have \( M=5 \) minibathces, +then each minibatch contains two data points. In particular we have \( B_1 = (\mathbf{x}_1,\mathbf{x}_2), \cdots, B_5 = (\mathbf{x}_9,\mathbf{x}_{10}) \). Note that if you choose \( M=1 \) you -have only a single batch with all datapoints and on the other extreme, +have only a single batch with all data points and on the other extreme, you may choose \( M=n \) resulting in a minibatch for each datapoint, i.e \( B_k = \mathbf{x}_k \).

    The idea is now to approximate the gradient by replacing the sum over -all datapoints with a sum over the datapoints in one the minibatches +all data points with a sum over the data points in one the minibatches picked at random in each gradient descent step $$ -\nabla_\beta +\nabla_{\beta} C(\mathbf{\beta}) = \sum_{i=1}^n \nabla_\beta c_i(\mathbf{x}_i, \mathbf{\beta}) \rightarrow \sum_{i \in B_k}^n \nabla_\beta c_i(\mathbf{x}_i, \mathbf{\beta}). @@ -747,8 +746,8 @@ Taking the gradient only on a subset of the data has two important benefits. First, it introduces randomness which decreases the chance that our opmization scheme gets stuck in a local minima. Second, if the size of the minibatches are small relative to the number of -datapoints (\( M < n \)), the computation of the gradient is much -cheaper since we sum over the datapoints in the k-th minibatch and not +datapoints (\( M < n \)), the computation of the gradient is much +cheaper since we sum over the datapoints in the \( k-th \) minibatch and not all \( n \) datapoints.

    @@ -779,7 +778,7 @@ number of epochs in such a way that it becomes very small after a reasonable time such that we do not move at all.

    -As an example, let \( e = 0,1,2,3,\cdots \) denote the current epoch and let \( t_0, t_1 > 0 \) be two fixed numbers. Furthermore, let \( t = e \cdot m + i \) where \( m \) is the number of minibatches and \( i=0,\cdots,m-1 \). Then the function $$\gamma_j(t; t_0, t_1) = \frac{t_0}{t+t_1} $$ goes to zero as the number of epochs gets large. I.e. we start with a step length \( \gamma_j (0; t_0, t_1) = t_0/t_1 \) which decays in time \( t \). +As an example, let \( e = 0,1,2,3,\cdots \) denote the current epoch and let \( t_0, t_1 > 0 \) be two fixed numbers. Furthermore, let \( t = e \cdot m + i \) where \( m \) is the number of minibatches and \( i=0,\cdots,m-1 \). Then the function $$\gamma_j(t; t_0, t_1) = \frac{t_0}{t+t_1} $$ goes to zero as the number of epochs gets large. I.e. we start with a step length \( \gamma_j (0; t_0, t_1) = t_0/t_1 \) which decays in time \( t \).

    In this way we can fix the number of epochs, compute \( \beta \) and diff --git a/doc/pub/Splines/html/Splines.html b/doc/pub/Splines/html/Splines.html index b7f5a50f9..eec927569 100644 --- a/doc/pub/Splines/html/Splines.html +++ b/doc/pub/Splines/html/Splines.html @@ -74,21 +74,21 @@ div { text-align: justify; text-justify: inter-word; } ('More on Steepest descent', 2, None, '___sec2'), ('The ideal', 2, None, '___sec3'), ('The sensitiveness of the gradient descent', 2, None, '___sec4'), - ('Gradient Descent Example', 2, None, '___sec5'), + ('Convex functions', 2, None, '___sec5'), + ('Convex function', 2, None, '___sec6'), + ('Conditions on convex functions', 2, None, '___sec7'), + ('More on convex functions', 2, None, '___sec8'), + ('Some simple problems', 2, None, '___sec9'), + ('Revisiting our first homework', 2, None, '___sec10'), + ('Gradient descent example', 2, None, '___sec11'), + ('The derivative of the cost/loss function', 2, None, '___sec12'), + ('The Hessian matrix', 2, None, '___sec13'), + ('Simple program', 2, None, '___sec14'), + ('Gradient Descent Example', 2, None, '___sec15'), ('And a corresponding example using _scikit-learn_', 2, None, - '___sec6'), - ('Convex functions', 2, None, '___sec7'), - ('Convex function', 2, None, '___sec8'), - ('Conditions on convex functions', 2, None, '___sec9'), - ('More on convex functions', 2, None, '___sec10'), - ('Some simple problems', 2, None, '___sec11'), - ('Revisiting our first homework', 2, None, '___sec12'), - ('Gradient descent example', 2, None, '___sec13'), - ('The derivative of the cost/loss function', 2, None, '___sec14'), - ('The Hessian matrix', 2, None, '___sec15'), - ('Simple program', 2, None, '___sec16'), + '___sec16'), ('Gradient descent and Ridge', 2, None, '___sec17'), ('Stochastic Gradient Descent', 2, None, '___sec18'), ('Computation of gradients', 2, None, '___sec19'), @@ -152,7 +152,7 @@ MathJax.Hub.Config({

    [2] Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University

    -

    Sep 20, 2018

    +

    Sep 21, 2018












    @@ -183,11 +183,13 @@ direction of the negative gradient \( -\nabla F(\mathbf{x}) \).

    It can be shown that if $$ -\mathbf{x}_{k+1} = \mathbf{x}_k - \gamma_k \nabla F(\mathbf{x}_k), \ \ \gamma_k > 0 +\mathbf{x}_{k+1} = \mathbf{x}_k - \gamma_k \nabla F(\mathbf{x}_k), $$ +with \( \gamma_k > 0 \). +

    -for \( \gamma_k \) small enough, then \( F(\mathbf{x}_{k+1}) \leq +For \( \gamma_k \) small enough, then \( F(\mathbf{x}_{k+1}) \leq F(\mathbf{x}_k) \). This means that for a sufficiently small \( \gamma_k \) we are always moving towards smaller function values, i.e a minimum. @@ -216,7 +218,7 @@ the learning rate within the context of Machine Learning.

    The ideal

    -Ideally the sequence \( \{ \mathbf{x}_k \}_{k=0} \) converges to a global +Ideally the sequence \( \{\mathbf{x}_k \}_{k=0} \) converges to a global minimum of the function \( F \). In general we do not know if we are in a global or local minimum. In the special case when \( F \) is a convex function, all local minima are also global minima, so in this case @@ -242,7 +244,8 @@ Note that the gradient is a function of \( \mathbf{x} =

    The sensitiveness of the gradient descent

    -GD is sensitive to the choice of learning rate \( \gamma_k \). This is due +The gradient descent method +is sensitive to the choice of learning rate \( \gamma_k \). This is due to the fact that we are only guaranteed that \( F(\mathbf{x}_{k+1}) \leq F(\mathbf{x}_k) \) for sufficiently small \( \gamma_k \). The problem is to determine an optimal learning rate. If the learning rate is chosen too @@ -255,12 +258,267 @@ randomness. One such method is that of Stochastic Gradient Descent (SGD), see below.

    -









    + -

    Gradient Descent Example

    +

    Convex functions

    -We revisit now our simple linear regression example with a linear polynomial. +Ideally we want our cost/loss function to be convex(concave). + +

    +First we give the definition of a convex set: A set \( C \) in +\( \mathbb{R}^n \) is said to be convex if, for all \( x \) and \( y \) in \( C \) and +all \( t \in (0,1) \) , the point \( (1 − t)x + ty \) also belongs to +C. Geometrically this means that every point on the line segment +connecting \( x \) and \( y \) is in \( C \) as discussed below. + +

    +The convex subsets of \( \mathbb{R} \) are the intervals of +\( \mathbb{R} \). Examples of convex sets of \( \mathbb{R}^2 \) are the +regular polygons (triangles, rectangles, pentagons, etc...). + +

    +









    + +

    Convex function

    + +

    +Convex function: Let \( X \subset \mathbb{R}^n \) be a convex set. Assume that the function \( f: X \rightarrow \mathbb{R} \) is continuous, then \( f \) is said to be convex if $$f(tx_1 + (1-t)x_2) \leq tf(x_1) + (1-t)f(x_2) $$ for all \( x_1, x_2 \in X \) and for all \( t \in [0,1] \). If \( \leq \) is replaced with a strict inequaltiy in the definition, we demand \( x_1 \neq x_2 \) and \( t\in(0,1) \) then \( f \) is said to be strictly convex. For a single variable function, convexity means that if you draw a straight line connecting \( f(x_1) \) and \( f(x_2) \), the value of the function on the interval \( [x_1,x_2] \) is always below the line as illustrated below. + +

    +









    + +

    Conditions on convex functions

    + +

    +In the following we state first and second-order conditions which +ensures convexity of a function \( f \). We write \( D_f \) to denote the +domain of \( f \), i.e the subset of \( R^n \) where \( f \) is defined. For more +details and proofs we refer to: S. Boyd and L. Vandenberghe. Convex Optimization. Cambridge University Press. + +

    +

    +First order condition. +

    +Suppose \( f \) is differentiable (i.e \( \nabla f(x) \) is well defined for +all \( x \) in the domain of \( f \)). Then \( f \) is convex if and only if \( D_f \) +is a convex set and $$f(y) \geq f(x) + \nabla f(x)^T (y-x) $$ holds +for all \( x,y \in D_f \). This condition means that for a convex function +the first order Taylor expansion (right hand side above) at any point +a global under estimator of the function. To convince yourself you can +make a drawing of \( f(x) = x^2+1 \) and draw the tangent line to \( f(x) \) and +note that it is always below the graph. +

    + + +

    +

    +Second order condition. +

    +Assume that \( f \) is twice +differentiable, i.e the Hessian matrix exists at each point in +\( D_f \). Then \( f \) is convex if and only if \( D_f \) is a convex set and its +Hessian is positive semi-definite for all \( x\in D_f \). For a +single-variable function this reduces to \( f''(x) \geq 0 \). Geometrically this means that \( f \) has nonnegative curvature +everywhere. +

    + + +

    +This condition is particularly useful since it gives us an procedure for determining if the function under consideration is convex, apart from using the definition. + +

    +









    + +

    More on convex functions

    + +

    +The next result is of great importance to us and the reason why we are +going on about convex functions. In machine learning we frequently +have to minimize a loss/cost function in order to find the best +parameters for the model we are considering. + +

    +Ideally we want the +global minimum (for high-dimensional models it is hard to know +if we have local or global minimum). However, if the cost/loss function +is convex the following result provides invaluable information: + +

    +

    +Any minimum is global for convex functions. +

    +Consider the problem of finding \( x \in \mathbb{R}^n \) such that \( f(x) \) +is minimal, where \( f \) is convex and differentiable. Then, any point +\( x^* \) that satisfies \( \nabla f(x^*) = 0 \) is a global minimum. +

    + + +

    +This result means that if we know that the cost/loss function is convex and we are able to find a minimum, we are guaranteed that it is a global minimum. + +

    +









    + +

    Some simple problems

    + +
      +
    1. Show that \( f(x)=x^2 \) is convex for \( x \in \mathbb{R} \) using the definition of convexity. Hint: If you re-write the definition, \( f \) is convex if the following holds for all \( x,y \in D_f \) and any \( \lambda \in [0,1] \) $\lambda f(x)+(1-\lambda)f(y)-f(\lambda x + (1-\lambda) y ) \geq 0$.
    2. +
    3. Using the second order condition show that the following functions are convex on the specified domain.
    4. + +
        +
      • \( f(x) = e^x \) is convex for \( x \in \mathbb{R} \).
      • +
      • \( g(x) = -\ln(x) \) is convex for \( x \in (0,\infty) \).
      • +
      + +
    5. Let \( f(x) = x^2 \) and \( g(x) = e^x \). Show that \( f(g(x)) \) and \( g(f(x)) \) is convex for \( x \in \mathbb{R} \). Also show that if \( f(x) \) is any convex function than \( h(x) = e^{f(x)} \) is convex.
    6. +
    7. A norm is any function that satisfy the following properties
    8. + +
        +
      • \( f(\alpha x) = |\alpha| f(x) \) for all \( \alpha \in \mathbb{R} \).
      • +
      • \( f(x+y) \leq f(x) + f(y) \)
      • +
      • \( f(x) \leq 0 \) for all \( x \in \mathbb{R}^n \) with equality if and only if \( x = 0 \)
      • +
      + +
    + +Using the definition of convexity, try to show that a function satisfying the properties above is convex (the third condition is not needed to show this). + +

    + + +

    Revisiting our first homework

    + +

    +We will use linear regression as a case study for the gradient descent +methods. Linear regression is a great test case for the gradient +descent methods discussed in the lectures since it has several +desirable properties such as: + +

      +
    1. An analytical solution (recall homework set 1).
    2. +
    3. The gradient can be computed analytically.
    4. +
    5. The cost function is convex which guarantees that gradient descent converges for small enough learning rates
    6. +
    + +We revisit the example from homework set 1 where we had +$$ +y_i = 5x_i^2 + 0.1\xi_i, \ i=1,\cdots,100 +$$ + +with \( x_i \in [0,1] \) chosen randomly with a uniform distribution. Additionally \( \xi_i \) represents stochastic noise chosen according to a normal distribution \( \cal {N}(0,1) \). +The linear regression model is given by +$$ +h_\beta(x) = \hat{y} = \beta_0 + \beta_1 x, +$$ + +such that +$$ +\hat{y}_i = \beta_0 + \beta_1 x_i. +$$ + +

    + + +

    Gradient descent example

    + +

    +Let \( \mathbf{y} = (y_1,\cdots,y_n)^T \), \( \mathbf{\hat{y}} = (\hat{y}_1,\cdots,\hat{y}_n)^T \) and \( \beta = (\beta_0, \beta_1)^T \) + +

    +It is convenient to write \( \mathbf{\hat{y}} = X\beta \) where \( X \in \mathbb{R}^{100 \times 2} \) is the design matrix given by +$$ +X \equiv \begin{bmatrix} +1 & x_1 \\ +\vdots & \vdots \\ +1 & x_{100} & \\ +\end{bmatrix}. +$$ + +The loss function is given by +$$ +C(\beta) = ||X\beta-\mathbf{y}||^2 = ||X\beta||^2 - 2 \mathbf{y}^T X\beta + ||\mathbf{y}||^2 = \sum_{i=1}^{100} (\beta_0 + \beta_1 x_i)^2 - 2 y_i (\beta_0 + \beta_1 x_i) + y_i^2 +$$ + +and we want to find \( \beta \) such that \( C(\beta) \) is minimized. + +

    +









    + +

    The derivative of the cost/loss function

    + +

    +Computing \( \partial C(\beta) / \partial \beta_0 \) and \( \partial C(\beta) / \partial \beta_1 \) we can show that the gradient can be written as +$$ +\nabla_{\beta} C(\beta) = (\partial C(\beta) / \partial \beta_0, \partial C(\beta) / \partial \beta_1)^T = 2\begin{bmatrix} \sum_{i=1}^{100} \left(\beta_0+\beta_1x_i-y_i\right) \\ +\sum_{i=1}^{100}\left( x_i (\beta_0+\beta_1x_i)-y_ix_i\right) \\ +\end{bmatrix} = 2X^T(X\beta - \mathbf{y}), +$$ + +where \( X \) is the design matrix defined above. + +

    +









    + +

    The Hessian matrix

    +The Hessian matrix of \( C(\beta) \) is given by +$$ +\hat{H} \equiv \begin{bmatrix} +\frac{\partial^2 C(\beta)}{\partial \beta_0^2} & \frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} \\ +\frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} & \frac{\partial^2 C(\beta)}{\partial \beta_1^2} & \\ +\end{bmatrix} = 2X^T X. +$$ + +This result implies that \( C(\beta) \) is a convex function since the matrix \( X^T X \) always is positive semi-definite. + +

    +









    + +

    Simple program

    + +

    +We can now write a program that minimizes \( C(\beta) \) using the gradient descent method with a constant learning rate \( \gamma \) according to +$$ +\beta_{k+1} = \beta_k - \gamma \nabla_\beta C(\beta_k), \ k=0,1,\cdots +$$ + +

    +We can use the expression we computed for the gradient and let use a +\( \beta_0 \) be chosen randomly and let \( \gamma = 0.001 \). Stop iterating +when \( ||\nabla_\beta C(\beta_k) || \leq \epsilon = 10^{-8} \). + +

    +And finally we can compare our solution for \( \beta \) with the analytic result given by +\( \beta= (X^TX)^{-1} X^T \mathbf{y} \). +

    + + +

    import numpy as np
    +
    +"""
    +The following setup is just a suggestion, feel free to write it the way you like.
    +"""
    +
    +#Setup problem described in the exercise
    +N  = 100 #Nr of datapoints
    +M  = 2 #Nr of features
    +x  = np.random.rand(N) #Uniformly generated x-values in [0,1]
    +y  = 5*x**2 + 0.1*np.random.randn(N)
    +X  = np.c_[np.ones(N),x] #Construct design matrix
    +
    +#Compute beta according to normal equations to compare with GD solution
    +Xt_X_inv = np.linalg.inv(np.dot(X.T,X))
    +Xt_y     = np.dot(X.transpose(),y)
    +beta_NE = np.dot(Xt_X_inv,Xt_y)
    +print(beta_NE)
    +
    +

    +









    + +

    Gradient Descent Example

    + +

    +Another simple example is here

    @@ -306,7 +564,7 @@ plt.show()











    -

    And a corresponding example using scikit-learn

    +

    And a corresponding example using scikit-learn

    @@ -330,265 +588,6 @@ sgdreg.fit(x,y.

    -

    Convex functions

    - -

    -Ideally we want our cost/loss function to be convex(concave). - -

    -First we give the definition of a convex set: A set \( C \) in -\( \mathbb{R}^n \) is said to be convex if, for all \( x \) and \( y \) in \( C \) and -all \( t \in (0,1) \) , the point \( (1 − t)x + ty \) also belongs to -C. Geometrically this means that every point on the line segment -connecting \( x \) and \( y \) is in \( C \) as discussed below. - -

    -The convex subsets of \( \mathbb{R} \) are the intervals of -\( \mathbb{R} \). Examples of convex sets of \( \mathbb{R}^2 \) are the -regular polygons (triangles, rectangles, pentagons, etc...). - -

    -









    - -

    Convex function

    - -

    -Convex function: Let \( X \subset \mathbb{R}^n \) be a convex set. Assume that the function \( f: X \rightarrow \mathbb{R} \) is continuous, then \( f \) is said to be convex if $$f(tx_1 + (1-t)x_2) \leq tf(x_1) + (1-t)f(x_2) $$ for all \( x_1, x_2 \in X \) and for all \( t \in [0,1] \). If \( \leq \) is replaced with a strict inequaltiy in the definition, we demand \( x_1 \neq x_2 \) and \( t\in(0,1) \) then \( f \) is said to be strictly convex. For a single variable function, convexity means that if you draw a straight line connecting \( f(x_1) \) and \( f(x_2) \), the value of the function on the interval \( [x_1,x_2] \) is always below the line as illustrated below. - -

    -









    - -

    Conditions on convex functions

    - -

    -In the following we state first and second-order conditions which -ensures convexity of a function \( f \). We write \( D_f \) to denote the -domain of \( f \), i.e the subset of \( R^n \) where \( f \) is defined. For more -details and proofs we refer to: S. Boyd and L. Vandenberghe. Convex Optimization. Cambridge University Press. - -

    -

    -First order condition. -

    -Suppose \( f \) is differentiable (i.e \( \nabla f(x) \) is well defined for -all \( x \) in the domain of \( f \)). Then \( f \) is convex if and only if \( D_f \) -is a convex set and $$f(y) \geq f(x) + \nabla f(x)^T (y-x) $$ holds -for all \( x,y \in D_f \). This condition means that for a convex function -the first order Taylor expansion (right hand side above) at any point -a global under estimator of the function. To convince yourself you can -make a drawing of f(x) = x^2+1 and draw the tangent line to \( f(x) \) and -note that it is always below the graph. -

    - - -

    -

    -Second order condition. -

    -Assume that \( f \) is twice -differentiable, i.e the Hessian matrix exists at each point in -\( D_f \). Then \( f \) is convex if and only if \( D_f \) is a convex set and its -Hessian is positive semi-definite for all \( x\in D_f \). For a -single-variable function this reduces to \( f''(x) \geq -0 \). Geometrically this means that \( f \) has nonnegative curvature -everywhere. -

    - - -

    -This condition is particularly useful since it gives us an procedure for determining if the function under consideration is convex, apart from using the definition. - -

    -









    - -

    More on convex functions

    - -

    -The next result is of great importance to us and the reason why we are -going on about convex functions. In machine learning we frequently -have to minimize a loss/cost function in order to find the best -parameters for the model we are considering. - -

    -Ideally we want the -global minimum (for high-dimensional models it is hard to know -if we have local or global minimum). However, if the cost/loss function -is convex the following result provides invaluable information: - -

    -

    -Any minimum is global for convex functions. -

    -Consider the problem of finding \( x \in \mathbb{R}^n \) such that \( f(x) \) -is minimal, where \( f \) is convex and differentiable. Then, any point -\( x^* \) that satisfies \( \nabla f(x^*) = 0 \) is a global minimum. -

    - - -

    -This result means that if we know that the cost/loss function is convex and we are able to find a minimum, we are guaranteed that it is a global minimum. - -

    -









    - -

    Some simple problems

    - -
      -
    1. Show that \( f(x)=x^2 \) is convex for \( x \in \mathbb{R} \) using the definition of convexity. Hint: If you re-write the definition, \( f \) is convex if the following holds for all \( x,y \in D_f \) and any \( \lambda \in [0,1] \) $\lambda f(x) + (1-\lambda)f(y) - f(\lambda x + (1-\lambda) y ) \geq 0. $
    2. -
    3. Using the second order condition show that the following functions are convex on the specified domain.
    4. - -
        -
      • \( f(x) = e^x \) is convex for \( x \in \mathbb{R} \).
      • -
      • \( g(x) = -\ln(x) \) is convex for \( x \in (0,\infty) \).
      • -
      - -
    5. Let \( f(x) = x^2 \) and \( g(x) = e^x \). Show that \( f(g(x)) \) and \( g(f(x)) \) is convex for \( x \in \mathbb{R} \). Also show that if \( f(x) \) is any convex function than \( h(x) = e^{f(x)} \) is convex.
    6. -
    7. A norm is any function that satisfy the following properties
    8. - -
        -
      • \( f(\alpha x) = |\alpha| f(x) \) for all \( \alpha \in \mathbb{R} \).
      • -
      • \( f(x+y) \leq f(x) + f(y) \)
      • -
      • \( f(x) \leq 0 \) for all \( x \in \mathbb{R}^n \) with equality if and only if \( x = 0 \)
      • -
      - -
    - -Using the definition of convexity, try to show that a function satisfying the properties above is convex (the third condition is not needed to show this). - -

    - - -

    Revisiting our first homework

    - -

    -We will use linear regression as a case study for the gradient descent -methods. Linear regression is a great test case for the gradient -descent methods discussed in the lectures since it has several -desirable properties such as: - -

      -
    1. An analytical solution (recall homework set 1).
    2. -
    3. The gradient can be computed analytically.
    4. -
    5. The cost function is convex which guarantees that gradient descent converges for small enough learning rates
    6. -
    - -We revisit the example from homework set 1 where we had -$$ -y_i = 5x_i^2 + 0.1\xi_i, \ i=1,\cdots,100 -$$ - -with \( x_i \in [0,1] \) chosen randomly with a uniform distribution. Additionally \( \xi_i \) represents stochastic noise chosen according to a normal distribution \( \cal {N}(0,1) \). -The linear regression model is given by -$$ -h_\beta(x) = \hat{y} = \beta_0 + \beta_1 x, -$$ - -such that -$$ -\hat{y}_i = \beta_0 + \beta_1 x_i. -$$ - -

    - - -

    Gradient descent example

    - -

    -Let \( \mathbf{y} = (y_1,\cdots,y_n)^T \), \( \mathbf{\hat{y}} = (\hat{y}_1,\cdots,\hat{y}_n)^T \) and \( \beta = (\beta_0, \beta_1)^T \) - -

    -t is convenient to write \( \mathbf{\hat{y}} = X\beta \) where \( X \in \mathbb{R}^{100 \times 2} \) is the design matrix given by -$$ -\begin{equation} -X \equiv \begin{bmatrix} -1 & x_1 \\ -\vdots & \vdots \\ -1 & x_{100} & \\ -\end{bmatrix}. -\label{_auto1} -\end{equation} -$$ - -The loss function is given by -$$ -C(\beta) = ||X\beta-\mathbf{y}||^2 = ||X\beta||^2 - 2 \mathbf{y}^T X\beta + ||\mathbf{y}||^2 = \sum_{i=1}^{100} (\beta_0 + \beta_1 x_i)^2 - 2 y_i (\beta_0 + \beta_1 x_i) + y_i^2 -$$ - -and we want to find \( \beta \) such that \( C(\beta) \) is minimized. - -

    -









    - -

    The derivative of the cost/loss function

    - -

    -Computing \( \partial C(\beta) / \partial \beta_0 \) and \( \partial C(\beta) / \partial \beta_1 \) we can show that the gradient can be written as -$$ -\nabla_\beta C(\beta) = (\partial C(\beta) / \partial \beta_0, \partial C(\beta) / \partial \beta_1)^T = 2\begin{bmatrix} \sum_{i=1}^{100} \left(\beta_0+\beta_1x_i-y_i\right) \\ -\sum_{i=1}^{100}\left( x_i (\beta_0+\beta_1x_i)-y_ix_i\right) \\ -\end{bmatrix} = 2X^T(X\beta - \mathbf{y}), -$$ - -where \( X \) is the design matrix defined above. - -

    -









    - -

    The Hessian matrix

    -The Hessian matrix of \( C(\beta) \) is given by -$$ -\hat{H} \equiv \begin{bmatrix} -\frac{\partial^2 C(\beta)}{\partial \beta_0^2} & \frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} \\ -\frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} & \frac{\partial^2 C(\beta)}{\partial \beta_1^2} & \\ -\end{bmatrix} = 2X^T X. -$$ - -This result implies that \( C(\beta) \) is a convex function since the matrix \( X^T X \) always is positive semi-definite. - -

    -









    - -

    Simple program

    - -

    -We can now write a program that minimizes \( C(\beta) \) using the gradient descent method with a constant learning rate \( \gamma \) according to -$$ -\beta_{k+1} = \beta_k - \gamma \nabla_\beta C(\beta_k), \ k=0,1,\cdots -$$ - -

    -We can use the expression we computed for the gradient and let use a -\( \beta_0 \) be chosen randomly and let \( \gamma = 0.001 \). Stop iterating -when \( ||\nabla_\beta C(\beta_k) || < \epsilon = 10^{-8} \). - -

    -And finally we can compare our solution for \( \beta \) with the analytic result given by -\( \beta= (X^TX)^{-1} X^T \mathbf{y} \). -

    - - -

    import numpy as np
    -
    -"""
    -The following setup is just a suggestion, feel free to write it the way you like.
    -"""
    -
    -#Setup problem described in the exercise
    -N  = 100 #Nr of datapoints
    -M  = 2 #Nr of features
    -x  = np.random.rand(N) #Uniformly generated x-values in [0,1]
    -y  = 5*x**2 + 0.1*np.random.randn(N)
    -X  = np.c_[np.ones(N),x] #Construct design matrix
    -
    -#Compute beta according to normal equations to compare with GD solution
    -Xt_X_inv = np.linalg.inv(np.dot(X.T,X))
    -Xt_y     = np.dot(X.transpose(),y)
    -beta_NE = np.dot(Xt_X_inv,Xt_y)
    -print(beta_NE)
    -
    -

    - -

    Gradient descent and Ridge

    @@ -655,7 +654,7 @@ the shortcomings of the Gradient descent method discussed above.

    The underlying idea of SGD comes from the observation that the cost function, which we want to minimize, can almost always be written as a -sum over \( n \) datapoints \( \{\mathbf{x}_i\}_{i=1}^n \), +sum over \( n \) data points \( \{\mathbf{x}_i\}_{i=1}^n \), $$ C(\mathbf{\beta}) = \sum_{i=1}^n c_i(\mathbf{x}_i, \mathbf{\beta}). @@ -677,7 +676,7 @@ $$

    Stochasticity/randomness is introduced by only taking the gradient on a subset of the data called minibatches. If there are \( n \) -datapoints and the size of each minibatch is \( M \), there will be \( n/M \) +data points and the size of each minibatch is \( M \), there will be \( n/M \) minibatches. We denote these minibatches by \( B_k \) where \( k=1,\cdots,n/M \). @@ -685,21 +684,21 @@ minibatches. We denote these minibatches by \( B_k \) where









    SGD example

    -As an example, suppose we have \( 10 \) datapoints \( ( \mathbf{x}_1, -\cdots, \mathbf{x}_{10} ) \) and we choose to have \( M=5 \) minibathces, -then each minibatch contains two datapoints. In particular we have +As an example, suppose we have \( 10 \) data points \( (\mathbf{x}_1,\cdots, \mathbf{x}_{10}) \) +and we choose to have \( M=5 \) minibathces, +then each minibatch contains two data points. In particular we have \( B_1 = (\mathbf{x}_1,\mathbf{x}_2), \cdots, B_5 = (\mathbf{x}_9,\mathbf{x}_{10}) \). Note that if you choose \( M=1 \) you -have only a single batch with all datapoints and on the other extreme, +have only a single batch with all data points and on the other extreme, you may choose \( M=n \) resulting in a minibatch for each datapoint, i.e \( B_k = \mathbf{x}_k \).

    The idea is now to approximate the gradient by replacing the sum over -all datapoints with a sum over the datapoints in one the minibatches +all data points with a sum over the data points in one the minibatches picked at random in each gradient descent step $$ -\nabla_\beta +\nabla_{\beta} C(\mathbf{\beta}) = \sum_{i=1}^n \nabla_\beta c_i(\mathbf{x}_i, \mathbf{\beta}) \rightarrow \sum_{i \in B_k}^n \nabla_\beta c_i(\mathbf{x}_i, \mathbf{\beta}). @@ -752,8 +751,8 @@ Taking the gradient only on a subset of the data has two important benefits. First, it introduces randomness which decreases the chance that our opmization scheme gets stuck in a local minima. Second, if the size of the minibatches are small relative to the number of -datapoints (\( M < n \)), the computation of the gradient is much -cheaper since we sum over the datapoints in the k-th minibatch and not +datapoints (\( M < n \)), the computation of the gradient is much +cheaper since we sum over the datapoints in the \( k-th \) minibatch and not all \( n \) datapoints.

    @@ -784,7 +783,7 @@ number of epochs in such a way that it becomes very small after a reasonable time such that we do not move at all.

    -As an example, let \( e = 0,1,2,3,\cdots \) denote the current epoch and let \( t_0, t_1 > 0 \) be two fixed numbers. Furthermore, let \( t = e \cdot m + i \) where \( m \) is the number of minibatches and \( i=0,\cdots,m-1 \). Then the function $$\gamma_j(t; t_0, t_1) = \frac{t_0}{t+t_1} $$ goes to zero as the number of epochs gets large. I.e. we start with a step length \( \gamma_j (0; t_0, t_1) = t_0/t_1 \) which decays in time \( t \). +As an example, let \( e = 0,1,2,3,\cdots \) denote the current epoch and let \( t_0, t_1 > 0 \) be two fixed numbers. Furthermore, let \( t = e \cdot m + i \) where \( m \) is the number of minibatches and \( i=0,\cdots,m-1 \). Then the function $$\gamma_j(t; t_0, t_1) = \frac{t_0}{t+t_1} $$ goes to zero as the number of epochs gets large. I.e. we start with a step length \( \gamma_j (0; t_0, t_1) = t_0/t_1 \) which decays in time \( t \).

    In this way we can fix the number of epochs, compute \( \beta \) and diff --git a/doc/pub/Splines/ipynb/Splines.ipynb b/doc/pub/Splines/ipynb/Splines.ipynb index 5328f5ba0..4f4ff5c1d 100644 --- a/doc/pub/Splines/ipynb/Splines.ipynb +++ b/doc/pub/Splines/ipynb/Splines.ipynb @@ -10,7 +10,7 @@ " \n", "**Morten Hjorth-Jensen**, Department of Physics, University of Oslo and Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University\n", "\n", - "Date: **Sep 20, 2018**\n", + "Date: **Sep 21, 2018**\n", "\n", "Copyright 1999-2018, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license\n", "\n", @@ -43,7 +43,7 @@ "metadata": {}, "source": [ "$$\n", - "\\mathbf{x}_{k+1} = \\mathbf{x}_k - \\gamma_k \\nabla F(\\mathbf{x}_k), \\ \\ \\gamma_k > 0\n", + "\\mathbf{x}_{k+1} = \\mathbf{x}_k - \\gamma_k \\nabla F(\\mathbf{x}_k),\n", "$$" ] }, @@ -51,7 +51,9 @@ "cell_type": "markdown", "metadata": {}, "source": [ - "for $\\gamma_k$ small enough, then $F(\\mathbf{x}_{k+1}) \\leq\n", + "with $\\gamma_k > 0$.\n", + "\n", + "For $\\gamma_k$ small enough, then $F(\\mathbf{x}_{k+1}) \\leq\n", "F(\\mathbf{x}_k)$. This means that for a sufficiently small $\\gamma_k$\n", "we are always moving towards smaller function values, i.e a minimum.\n", "\n", @@ -83,7 +85,7 @@ "\n", "## The ideal\n", "\n", - "Ideally the sequence $\\{ \\mathbf{x}_k \\}_{k=0}$ converges to a global\n", + "Ideally the sequence $\\{\\mathbf{x}_k \\}_{k=0}$ converges to a global\n", "minimum of the function $F$. In general we do not know if we are in a\n", "global or local minimum. In the special case when $F$ is a convex\n", "function, all local minima are also global minima, so in this case\n", @@ -105,7 +107,8 @@ "\n", "## The sensitiveness of the gradient descent\n", "\n", - "GD is sensitive to the choice of learning rate $\\gamma_k$. This is due\n", + "The gradient descent method \n", + "is sensitive to the choice of learning rate $\\gamma_k$. This is due\n", "to the fact that we are only guaranteed that $F(\\mathbf{x}_{k+1}) \\leq\n", "F(\\mathbf{x}_k)$ for sufficiently small $\\gamma_k$. The problem is to\n", "determine an optimal learning rate. If the learning rate is chosen too\n", @@ -116,98 +119,7 @@ "randomness. One such method is that of Stochastic Gradient Descent\n", "(SGD), see below.\n", "\n", - "## Gradient Descent Example\n", "\n", - "We revisit now our simple linear regression example with a linear polynomial." - ] - }, - { - "cell_type": "code", - "execution_count": 1, - "metadata": { - "collapsed": false - }, - "outputs": [], - "source": [ - "%matplotlib inline\n", - "\n", - "\n", - "# Importing various packages\n", - "from random import random, seed\n", - "import numpy as np\n", - "import matplotlib.pyplot as plt\n", - "from mpl_toolkits.mplot3d import Axes3D\n", - "from matplotlib import cm\n", - "from matplotlib.ticker import LinearLocator, FormatStrFormatter\n", - "import sys\n", - "\n", - "x = 2*np.random.rand(100,1)\n", - "y = 4+3*x+np.random.randn(100,1)\n", - "\n", - "xb = np.c_[np.ones((100,1)), x]\n", - "beta_linreg = np.linalg.inv(xb.T.dot(xb)).dot(xb.T).dot(y)\n", - "print(beta_linreg)\n", - "beta = np.random.randn(2,1)\n", - "\n", - "eta = 0.1\n", - "Niterations = 1000\n", - "m = 100\n", - "\n", - "for iter in range(Niterations):\n", - " gradients = 2.0/m*xb.T.dot(xb.dot(beta)-y)\n", - " beta -= eta*gradients\n", - "\n", - "print(beta)\n", - "xnew = np.array([[0],[2]])\n", - "xbnew = np.c_[np.ones((2,1)), xnew]\n", - "ypredict = xbnew.dot(beta)\n", - "ypredict2 = xbnew.dot(beta_linreg)\n", - "plt.plot(xnew, ypredict, \"r-\")\n", - "plt.plot(xnew, ypredict2, \"b-\")\n", - "plt.plot(x, y ,'ro')\n", - "plt.axis([0,2.0,0, 15.0])\n", - "plt.xlabel(r'$x$')\n", - "plt.ylabel(r'$y$')\n", - "plt.title(r'Gradient descent example')\n", - "plt.show()" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "## And a corresponding example using **scikit-learn**" - ] - }, - { - "cell_type": "code", - "execution_count": 2, - "metadata": { - "collapsed": false - }, - "outputs": [], - "source": [ - "# Importing various packages\n", - "from random import random, seed\n", - "import numpy as np\n", - "import matplotlib.pyplot as plt\n", - "from sklearn.linear_model import SGDRegressor\n", - "\n", - "x = 2*np.random.rand(100,1)\n", - "y = 4+3*x+np.random.randn(100,1)\n", - "\n", - "xb = np.c_[np.ones((100,1)), x]\n", - "beta_linreg = np.linalg.inv(xb.T.dot(xb)).dot(xb.T).dot(y)\n", - "print(beta_linreg)\n", - "sgdreg = SGDRegressor(n_iter = 50, penalty=None, eta0=0.1)\n", - "sgdreg.fit(x,y.ravel())\n", - "print(sgdreg.intercept_, sgdreg.coef_)" - ] - }, - { - "cell_type": "markdown", - "metadata": {}, - "source": [ "\n", "## Convex functions\n", "\n", @@ -242,7 +154,7 @@ "for all $x,y \\in D_f$. This condition means that for a convex function\n", "the first order Taylor expansion (right hand side above) at any point\n", "a global under estimator of the function. To convince yourself you can\n", - "make a drawing of f(x) = x^2+1 and draw the tangent line to $f(x)$ and\n", + "make a drawing of $f(x) = x^2+1$ and draw the tangent line to $f(x)$ and\n", "note that it is always below the graph.\n", "\n", "\n", @@ -253,8 +165,7 @@ "differentiable, i.e the Hessian matrix exists at each point in\n", "$D_f$. Then $f$ is convex if and only if $D_f$ is a convex set and its\n", "Hessian is positive semi-definite for all $x\\in D_f$. For a\n", - "single-variable function this reduces to $f''(x) \\geq\n", - "0$. Geometrically this means that $f$ has nonnegative curvature\n", + "single-variable function this reduces to $f''(x) \\geq 0$. Geometrically this means that $f$ has nonnegative curvature\n", "everywhere.\n", "\n", "\n", @@ -285,7 +196,7 @@ "\n", "## Some simple problems\n", "\n", - "1. Show that $f(x)=x^2$ is convex for $x \\in \\mathbb{R}$ using the definition of convexity. Hint: If you re-write the definition, $f$ is convex if the following holds for all $x,y \\in D_f$ and any $\\lambda \\in [0,1] $ $\\lambda f(x) + (1-\\lambda)f(y) - f(\\lambda x + (1-\\lambda) y ) \\geq 0. $\n", + "1. Show that $f(x)=x^2$ is convex for $x \\in \\mathbb{R}$ using the definition of convexity. Hint: If you re-write the definition, $f$ is convex if the following holds for all $x,y \\in D_f$ and any $\\lambda \\in [0,1]$ $\\lambda f(x)+(1-\\lambda)f(y)-f(\\lambda x + (1-\\lambda) y ) \\geq 0$.\n", "\n", "2. Using the second order condition show that the following functions are convex on the specified domain.\n", "\n", @@ -375,25 +286,19 @@ "\n", "Let $\\mathbf{y} = (y_1,\\cdots,y_n)^T$, $\\mathbf{\\hat{y}} = (\\hat{y}_1,\\cdots,\\hat{y}_n)^T$ and $\\beta = (\\beta_0, \\beta_1)^T$\n", "\n", - "t is convenient to write $\\mathbf{\\hat{y}} = X\\beta$ where $X \\in \\mathbb{R}^{100 \\times 2} $ is the design matrix given by" + "It is convenient to write $\\mathbf{\\hat{y}} = X\\beta$ where $X \\in \\mathbb{R}^{100 \\times 2} $ is the design matrix given by" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ - "\n", - "

    \n", - "\n", "$$\n", - "\\begin{equation}\n", "X \\equiv \\begin{bmatrix}\n", "1 & x_1 \\\\\n", "\\vdots & \\vdots \\\\\n", "1 & x_{100} & \\\\\n", "\\end{bmatrix}.\n", - "\\label{_auto1} \\tag{1}\n", - "\\end{equation}\n", "$$" ] }, @@ -429,7 +334,7 @@ "metadata": {}, "source": [ "$$\n", - "\\nabla_\\beta C(\\beta) = (\\partial C(\\beta) / \\partial \\beta_0, \\partial C(\\beta) / \\partial \\beta_1)^T = 2\\begin{bmatrix} \\sum_{i=1}^{100} \\left(\\beta_0+\\beta_1x_i-y_i\\right) \\\\\n", + "\\nabla_{\\beta} C(\\beta) = (\\partial C(\\beta) / \\partial \\beta_0, \\partial C(\\beta) / \\partial \\beta_1)^T = 2\\begin{bmatrix} \\sum_{i=1}^{100} \\left(\\beta_0+\\beta_1x_i-y_i\\right) \\\\\n", "\\sum_{i=1}^{100}\\left( x_i (\\beta_0+\\beta_1x_i)-y_ix_i\\right) \\\\\n", "\\end{bmatrix} = 2X^T(X\\beta - \\mathbf{y}),\n", "$$" @@ -483,7 +388,7 @@ "source": [ "We can use the expression we computed for the gradient and let use a\n", "$\\beta_0$ be chosen randomly and let $\\gamma = 0.001$. Stop iterating\n", - "when $||\\nabla_\\beta C(\\beta_k) || < \\epsilon = 10^{-8}$. \n", + "when $||\\nabla_\\beta C(\\beta_k) || \\leq \\epsilon = 10^{-8}$. \n", "\n", "And finally we can compare our solution for $\\beta$ with the analytic result given by \n", "$\\beta= (X^TX)^{-1} X^T \\mathbf{y}$." @@ -491,7 +396,7 @@ }, { "cell_type": "code", - "execution_count": 3, + "execution_count": 1, "metadata": { "collapsed": false }, @@ -517,6 +422,98 @@ "print(beta_NE)" ] }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "## Gradient Descent Example\n", + "\n", + "Another simple example is here" + ] + }, + { + "cell_type": "code", + "execution_count": 2, + "metadata": { + "collapsed": false + }, + "outputs": [], + "source": [ + "%matplotlib inline\n", + "\n", + "\n", + "# Importing various packages\n", + "from random import random, seed\n", + "import numpy as np\n", + "import matplotlib.pyplot as plt\n", + "from mpl_toolkits.mplot3d import Axes3D\n", + "from matplotlib import cm\n", + "from matplotlib.ticker import LinearLocator, FormatStrFormatter\n", + "import sys\n", + "\n", + "x = 2*np.random.rand(100,1)\n", + "y = 4+3*x+np.random.randn(100,1)\n", + "\n", + "xb = np.c_[np.ones((100,1)), x]\n", + "beta_linreg = np.linalg.inv(xb.T.dot(xb)).dot(xb.T).dot(y)\n", + "print(beta_linreg)\n", + "beta = np.random.randn(2,1)\n", + "\n", + "eta = 0.1\n", + "Niterations = 1000\n", + "m = 100\n", + "\n", + "for iter in range(Niterations):\n", + " gradients = 2.0/m*xb.T.dot(xb.dot(beta)-y)\n", + " beta -= eta*gradients\n", + "\n", + "print(beta)\n", + "xnew = np.array([[0],[2]])\n", + "xbnew = np.c_[np.ones((2,1)), xnew]\n", + "ypredict = xbnew.dot(beta)\n", + "ypredict2 = xbnew.dot(beta_linreg)\n", + "plt.plot(xnew, ypredict, \"r-\")\n", + "plt.plot(xnew, ypredict2, \"b-\")\n", + "plt.plot(x, y ,'ro')\n", + "plt.axis([0,2.0,0, 15.0])\n", + "plt.xlabel(r'$x$')\n", + "plt.ylabel(r'$y$')\n", + "plt.title(r'Gradient descent example')\n", + "plt.show()" + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "## And a corresponding example using **scikit-learn**" + ] + }, + { + "cell_type": "code", + "execution_count": 3, + "metadata": { + "collapsed": false + }, + "outputs": [], + "source": [ + "# Importing various packages\n", + "from random import random, seed\n", + "import numpy as np\n", + "import matplotlib.pyplot as plt\n", + "from sklearn.linear_model import SGDRegressor\n", + "\n", + "x = 2*np.random.rand(100,1)\n", + "y = 4+3*x+np.random.randn(100,1)\n", + "\n", + "xb = np.c_[np.ones((100,1)), x]\n", + "beta_linreg = np.linalg.inv(xb.T.dot(xb)).dot(xb.T).dot(y)\n", + "print(beta_linreg)\n", + "sgdreg = SGDRegressor(n_iter = 50, penalty=None, eta0=0.1)\n", + "sgdreg.fit(x,y.ravel())\n", + "print(sgdreg.intercept_, sgdreg.coef_)" + ] + }, { "cell_type": "markdown", "metadata": {}, @@ -624,7 +621,7 @@ "\n", "The underlying idea of SGD comes from the observation that the cost\n", "function, which we want to minimize, can almost always be written as a\n", - "sum over $n$ datapoints $\\{\\mathbf{x}_i\\}_{i=1}^n$," + "sum over $n$ data points $\\{\\mathbf{x}_i\\}_{i=1}^n$," ] }, { @@ -663,22 +660,22 @@ "source": [ "Stochasticity/randomness is introduced by only taking the\n", "gradient on a subset of the data called minibatches. If there are $n$\n", - "datapoints and the size of each minibatch is $M$, there will be $n/M$\n", + "data points and the size of each minibatch is $M$, there will be $n/M$\n", "minibatches. We denote these minibatches by $B_k$ where\n", "$k=1,\\cdots,n/M$.\n", "\n", "## SGD example\n", - "As an example, suppose we have $10$ datapoints $( \\mathbf{x}_1,\n", - "\\cdots, \\mathbf{x}_{10} )$ and we choose to have $M=5$ minibathces,\n", - "then each minibatch contains two datapoints. In particular we have\n", + "As an example, suppose we have $10$ data points $(\\mathbf{x}_1,\\cdots, \\mathbf{x}_{10})$ \n", + "and we choose to have $M=5$ minibathces,\n", + "then each minibatch contains two data points. In particular we have\n", "$B_1 = (\\mathbf{x}_1,\\mathbf{x}_2), \\cdots, B_5 =\n", "(\\mathbf{x}_9,\\mathbf{x}_{10})$. Note that if you choose $M=1$ you\n", - "have only a single batch with all datapoints and on the other extreme,\n", + "have only a single batch with all data points and on the other extreme,\n", "you may choose $M=n$ resulting in a minibatch for each datapoint, i.e\n", "$B_k = \\mathbf{x}_k$.\n", "\n", "The idea is now to approximate the gradient by replacing the sum over\n", - "all datapoints with a sum over the datapoints in one the minibatches\n", + "all data points with a sum over the data points in one the minibatches\n", "picked at random in each gradient descent step" ] }, @@ -687,7 +684,7 @@ "metadata": {}, "source": [ "$$\n", - "\\nabla_\\beta\n", + "\\nabla_{\\beta}\n", "C(\\mathbf{\\beta}) = \\sum_{i=1}^n \\nabla_\\beta c_i(\\mathbf{x}_i,\n", "\\mathbf{\\beta}) \\rightarrow \\sum_{i \\in B_k}^n \\nabla_\\beta\n", "c_i(\\mathbf{x}_i, \\mathbf{\\beta}).\n", @@ -758,8 +755,8 @@ "benefits. First, it introduces randomness which decreases the chance\n", "that our opmization scheme gets stuck in a local minima. Second, if\n", "the size of the minibatches are small relative to the number of\n", - "datapoints ($M < n$), the computation of the gradient is much\n", - "cheaper since we sum over the datapoints in the k-th minibatch and not\n", + "datapoints ($M < n$), the computation of the gradient is much\n", + "cheaper since we sum over the datapoints in the $k-th$ minibatch and not\n", "all $n$ datapoints.\n", "\n", "## When do we stop?\n", @@ -781,7 +778,7 @@ "number of epochs in such a way that it becomes very small after a\n", "reasonable time such that we do not move at all.\n", "\n", - "As an example, let $e = 0,1,2,3,\\cdots$ denote the current epoch and let $t_0, t_1 > 0$ be two fixed numbers. Furthermore, let $t = e \\cdot m + i$ where $m$ is the number of minibatches and $i=0,\\cdots,m-1$. Then the function $$\\gamma_j(t; t_0, t_1) = \\frac{t_0}{t+t_1} $$ goes to zero as the number of epochs gets large. I.e. we start with a step length $\\gamma_j (0; t_0, t_1) = t_0/t_1$ which decays in *time* $t$.\n", + "As an example, let $e = 0,1,2,3,\\cdots$ denote the current epoch and let $t_0, t_1 > 0$ be two fixed numbers. Furthermore, let $t = e \\cdot m + i$ where $m$ is the number of minibatches and $i=0,\\cdots,m-1$. Then the function $$\\gamma_j(t; t_0, t_1) = \\frac{t_0}{t+t_1} $$ goes to zero as the number of epochs gets large. I.e. we start with a step length $\\gamma_j (0; t_0, t_1) = t_0/t_1$ which decays in *time* $t$.\n", "\n", "In this way we can fix the number of epochs, compute $\\beta$ and\n", "evaluate the cost function at the end. Repeating the computation will\n", diff --git a/doc/pub/Splines/ipynb/ipynb-Splines-src.tar.gz b/doc/pub/Splines/ipynb/ipynb-Splines-src.tar.gz index a1b16d968..34845ba2e 100644 Binary files a/doc/pub/Splines/ipynb/ipynb-Splines-src.tar.gz and b/doc/pub/Splines/ipynb/ipynb-Splines-src.tar.gz differ diff --git a/doc/pub/Splines/pdf/Splines-minted.pdf b/doc/pub/Splines/pdf/Splines-minted.pdf index c5ac6574a..b63252b15 100644 Binary files a/doc/pub/Splines/pdf/Splines-minted.pdf and b/doc/pub/Splines/pdf/Splines-minted.pdf differ diff --git a/doc/src/Splines/Splines.do.txt b/doc/src/Splines/Splines.do.txt index a1465d76d..950f1cad9 100644 --- a/doc/src/Splines/Splines.do.txt +++ b/doc/src/Splines/Splines.do.txt @@ -26,11 +26,12 @@ direction of the negative gradient $-\nabla F(\mathbf{x})$. It can be shown that if !bt \[ -\mathbf{x}_{k+1} = \mathbf{x}_k - \gamma_k \nabla F(\mathbf{x}_k), \ \ \gamma_k > 0 +\mathbf{x}_{k+1} = \mathbf{x}_k - \gamma_k \nabla F(\mathbf{x}_k), \] !et +with $\gamma_k > 0$. -for $\gamma_k$ small enough, then $F(\mathbf{x}_{k+1}) \leq +For $\gamma_k$ small enough, then $F(\mathbf{x}_{k+1}) \leq F(\mathbf{x}_k)$. This means that for a sufficiently small $\gamma_k$ we are always moving towards smaller function values, i.e a minimum. @@ -54,7 +55,7 @@ the learning rate within the context of Machine Learning. !split ===== The ideal ===== -Ideally the sequence $\{ \mathbf{x}_k \}_{k=0}$ converges to a global +Ideally the sequence $\{\mathbf{x}_k \}_{k=0}$ converges to a global minimum of the function $F$. In general we do not know if we are in a global or local minimum. In the special case when $F$ is a convex function, all local minima are also global minima, so in this case @@ -76,7 +77,8 @@ Note that the gradient is a function of $\mathbf{x} = !split ===== The sensitiveness of the gradient descent ===== -GD is sensitive to the choice of learning rate $\gamma_k$. This is due +The gradient descent method +is sensitive to the choice of learning rate $\gamma_k$. This is due to the fact that we are only guaranteed that $F(\mathbf{x}_{k+1}) \leq F(\mathbf{x}_k)$ for sufficiently small $\gamma_k$. The problem is to determine an optimal learning rate. If the learning rate is chosen too @@ -87,10 +89,217 @@ Many of these shortcomings can be alleviated by introducing randomness. One such method is that of Stochastic Gradient Descent (SGD), see below. + +!split +===== Convex functions ===== + +Ideally we want our cost/loss function to be convex(concave). + +First we give the definition of a convex set: A set $C$ in +$\mathbb{R}^n$ is said to be convex if, for all $x$ and $y$ in $C$ and +all $t \in (0,1)$ , the point $(1 − t)x + ty$ also belongs to +C. Geometrically this means that every point on the line segment +connecting $x$ and $y$ is in $C$ as discussed below. + +The convex subsets of $\mathbb{R}$ are the intervals of +$\mathbb{R}$. Examples of convex sets of $\mathbb{R}^2$ are the +regular polygons (triangles, rectangles, pentagons, etc...). + +!split +===== Convex function ===== + +_Convex function_: Let $X \subset \mathbb{R}^n$ be a convex set. Assume that the function $f: X \rightarrow \mathbb{R}$ is continuous, then $f$ is said to be convex if $$f(tx_1 + (1-t)x_2) \leq tf(x_1) + (1-t)f(x_2) $$ for all $x_1, x_2 \in X$ and for all $t \in [0,1]$. If $\leq$ is replaced with a strict inequaltiy in the definition, we demand $x_1 \neq x_2$ and $t\in(0,1)$ then $f$ is said to be strictly convex. For a single variable function, convexity means that if you draw a straight line connecting $f(x_1)$ and $f(x_2)$, the value of the function on the interval $[x_1,x_2]$ is always below the line as illustrated below. + +!split +===== Conditions on convex functions ===== + +In the following we state first and second-order conditions which +ensures convexity of a function $f$. We write $D_f$ to denote the +domain of $f$, i.e the subset of $R^n$ where $f$ is defined. For more +details and proofs we refer to: "S. Boyd and L. Vandenberghe. Convex Optimization. Cambridge University Press":"http://stanford.edu/boyd/cvxbook/, 2004". + +!bblock First order condition +Suppose $f$ is differentiable (i.e $\nabla f(x)$ is well defined for +all $x$ in the domain of $f$). Then $f$ is convex if and only if $D_f$ +is a convex set and $$f(y) \geq f(x) + \nabla f(x)^T (y-x) $$ holds +for all $x,y \in D_f$. This condition means that for a convex function +the first order Taylor expansion (right hand side above) at any point +a global under estimator of the function. To convince yourself you can +make a drawing of $f(x) = x^2+1$ and draw the tangent line to $f(x)$ and +note that it is always below the graph. +!eblock + +!bblock Second order condition +Assume that $f$ is twice +differentiable, i.e the Hessian matrix exists at each point in +$D_f$. Then $f$ is convex if and only if $D_f$ is a convex set and its +Hessian is positive semi-definite for all $x\in D_f$. For a +single-variable function this reduces to $f''(x) \geq 0$. Geometrically this means that $f$ has nonnegative curvature +everywhere. +!eblock + +This condition is particularly useful since it gives us an procedure for determining if the function under consideration is convex, apart from using the definition. + +!split +===== More on convex functions ===== + +The next result is of great importance to us and the reason why we are +going on about convex functions. In machine learning we frequently +have to minimize a loss/cost function in order to find the best +parameters for the model we are considering. + +Ideally we want the +global minimum (for high-dimensional models it is hard to know +if we have local or global minimum). However, if the cost/loss function +is convex the following result provides invaluable information: + +!bblock Any minimum is global for convex functions +Consider the problem of finding $x \in \mathbb{R}^n$ such that $f(x)$ +is minimal, where $f$ is convex and differentiable. Then, any point +$x^*$ that satisfies $\nabla f(x^*) = 0$ is a global minimum. +!eblock + +This result means that if we know that the cost/loss function is convex and we are able to find a minimum, we are guaranteed that it is a global minimum. + +!split +===== Some simple problems ===== + +o Show that $f(x)=x^2$ is convex for $x \in \mathbb{R}$ using the definition of convexity. Hint: If you re-write the definition, $f$ is convex if the following holds for all $x,y \in D_f$ and any $\lambda \in [0,1]$ $\lambda f(x)+(1-\lambda)f(y)-f(\lambda x + (1-\lambda) y ) \geq 0$. + +o Using the second order condition show that the following functions are convex on the specified domain. + * $f(x) = e^x$ is convex for $x \in \mathbb{R}$. + * $g(x) = -\ln(x)$ is convex for $x \in (0,\infty)$. +o Let $f(x) = x^2$ and $g(x) = e^x$. Show that $f(g(x))$ and $g(f(x))$ is convex for $x \in \mathbb{R}$. Also show that if $f(x)$ is any convex function than $h(x) = e^{f(x)}$ is convex. + +o A norm is any function that satisfy the following properties + * $f(\alpha x) = |\alpha| f(x)$ for all $\alpha \in \mathbb{R}$. + * $f(x+y) \leq f(x) + f(y)$ + * $f(x) \leq 0$ for all $x \in \mathbb{R}^n$ with equality if and only if $x = 0$ + +Using the definition of convexity, try to show that a function satisfying the properties above is convex (the third condition is not needed to show this). + +!split +===== Revisiting our first homework ===== + +We will use linear regression as a case study for the gradient descent +methods. Linear regression is a great test case for the gradient +descent methods discussed in the lectures since it has several +desirable properties such as: + +o An analytical solution (recall homework set 1). +o The gradient can be computed analytically. +o The cost function is convex which guarantees that gradient descent converges for small enough learning rates + +We revisit the example from homework set 1 where we had +!bt +\[ +y_i = 5x_i^2 + 0.1\xi_i, \ i=1,\cdots,100 +\] +!et +with $x_i \in [0,1] $ chosen randomly with a uniform distribution. Additionally $\xi_i$ represents stochastic noise chosen according to a normal distribution $\cal {N}(0,1)$. +The linear regression model is given by +!bt +\[ +h_\beta(x) = \hat{y} = \beta_0 + \beta_1 x, +\] +!et +such that +!bt +\[ +\hat{y}_i = \beta_0 + \beta_1 x_i. +\] +!et + +!split +===== Gradient descent example ===== + +Let $\mathbf{y} = (y_1,\cdots,y_n)^T$, $\mathbf{\hat{y}} = (\hat{y}_1,\cdots,\hat{y}_n)^T$ and $\beta = (\beta_0, \beta_1)^T$ + +It is convenient to write $\mathbf{\hat{y}} = X\beta$ where $X \in \mathbb{R}^{100 \times 2} $ is the design matrix given by +!bt +\[ +X \equiv \begin{bmatrix} +1 & x_1 \\ +\vdots & \vdots \\ +1 & x_{100} & \\ +\end{bmatrix}. +\] +!et +The loss function is given by +!bt +\[ +C(\beta) = ||X\beta-\mathbf{y}||^2 = ||X\beta||^2 - 2 \mathbf{y}^T X\beta + ||\mathbf{y}||^2 = \sum_{i=1}^{100} (\beta_0 + \beta_1 x_i)^2 - 2 y_i (\beta_0 + \beta_1 x_i) + y_i^2 +\] +!et +and we want to find $\beta$ such that $C(\beta)$ is minimized. + +!split +===== The derivative of the cost/loss function ===== + +Computing $\partial C(\beta) / \partial \beta_0$ and $\partial C(\beta) / \partial \beta_1$ we can show that the gradient can be written as +!bt +\[ +\nabla_{\beta} C(\beta) = (\partial C(\beta) / \partial \beta_0, \partial C(\beta) / \partial \beta_1)^T = 2\begin{bmatrix} \sum_{i=1}^{100} \left(\beta_0+\beta_1x_i-y_i\right) \\ +\sum_{i=1}^{100}\left( x_i (\beta_0+\beta_1x_i)-y_ix_i\right) \\ +\end{bmatrix} = 2X^T(X\beta - \mathbf{y}), +\] +!et +where $X$ is the design matrix defined above. + +!split +===== The Hessian matrix ===== +The Hessian matrix of $C(\beta)$ is given by +!bt +\[ +\hat{H} \equiv \begin{bmatrix} +\frac{\partial^2 C(\beta)}{\partial \beta_0^2} & \frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} \\ +\frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} & \frac{\partial^2 C(\beta)}{\partial \beta_1^2} & \\ +\end{bmatrix} = 2X^T X. +\] +!et +This result implies that $C(\beta)$ is a convex function since the matrix $X^T X$ always is positive semi-definite. + +!split +===== Simple program ===== + +We can now write a program that minimizes $C(\beta)$ using the gradient descent method with a constant learning rate $\gamma$ according to +!bt +\[ +\beta_{k+1} = \beta_k - \gamma \nabla_\beta C(\beta_k), \ k=0,1,\cdots +\] +!et + +We can use the expression we computed for the gradient and let use a +$\beta_0$ be chosen randomly and let $\gamma = 0.001$. Stop iterating +when $||\nabla_\beta C(\beta_k) || \leq \epsilon = 10^{-8}$. + +And finally we can compare our solution for $\beta$ with the analytic result given by +$\beta= (X^TX)^{-1} X^T \mathbf{y}$. +!bc pycod +import numpy as np + +""" +The following setup is just a suggestion, feel free to write it the way you like. +""" + +#Setup problem described in the exercise +N = 100 #Nr of datapoints +M = 2 #Nr of features +x = np.random.rand(N) #Uniformly generated x-values in [0,1] +y = 5*x**2 + 0.1*np.random.randn(N) +X = np.c_[np.ones(N),x] #Construct design matrix + +#Compute beta according to normal equations to compare with GD solution +Xt_X_inv = np.linalg.inv(np.dot(X.T,X)) +Xt_y = np.dot(X.transpose(),y) +beta_NE = np.dot(Xt_X_inv,Xt_y) +print(beta_NE) +!ec + !split ===== Gradient Descent Example ===== -We revisit now our simple linear regression example with a linear polynomial. +Another simple example is here !bc pycod # Importing various packages @@ -156,214 +365,7 @@ print(sgdreg.intercept_, sgdreg.coef_) !ec -!split -===== Convex functions ===== -Ideally we want our cost/loss function to be convex(concave). - -First we give the definition of a convex set: A set $C$ in -$\mathbb{R}^n$ is said to be convex if, for all $x$ and $y$ in $C$ and -all $t \in (0,1)$ , the point $(1 − t)x + ty$ also belongs to -C. Geometrically this means that every point on the line segment -connecting $x$ and $y$ is in $C$ as discussed below. - -The convex subsets of $\mathbb{R}$ are the intervals of -$\mathbb{R}$. Examples of convex sets of $\mathbb{R}^2$ are the -regular polygons (triangles, rectangles, pentagons, etc...). - -!split -===== Convex function ===== - -_Convex function_: Let $X \subset \mathbb{R}^n$ be a convex set. Assume that the function $f: X \rightarrow \mathbb{R}$ is continuous, then $f$ is said to be convex if $$f(tx_1 + (1-t)x_2) \leq tf(x_1) + (1-t)f(x_2) $$ for all $x_1, x_2 \in X$ and for all $t \in [0,1]$. If $\leq$ is replaced with a strict inequaltiy in the definition, we demand $x_1 \neq x_2$ and $t\in(0,1)$ then $f$ is said to be strictly convex. For a single variable function, convexity means that if you draw a straight line connecting $f(x_1)$ and $f(x_2)$, the value of the function on the interval $[x_1,x_2]$ is always below the line as illustrated below. - -!split -===== Conditions on convex functions ===== - -In the following we state first and second-order conditions which -ensures convexity of a function $f$. We write $D_f$ to denote the -domain of $f$, i.e the subset of $R^n$ where $f$ is defined. For more -details and proofs we refer to: "S. Boyd and L. Vandenberghe. Convex Optimization. Cambridge University Press":"http://stanford.edu/boyd/cvxbook/, 2004". - -!bblock First order condition -Suppose $f$ is differentiable (i.e $\nabla f(x)$ is well defined for -all $x$ in the domain of $f$). Then $f$ is convex if and only if $D_f$ -is a convex set and $$f(y) \geq f(x) + \nabla f(x)^T (y-x) $$ holds -for all $x,y \in D_f$. This condition means that for a convex function -the first order Taylor expansion (right hand side above) at any point -a global under estimator of the function. To convince yourself you can -make a drawing of f(x) = x^2+1 and draw the tangent line to $f(x)$ and -note that it is always below the graph. -!eblock - -!bblock Second order condition -Assume that $f$ is twice -differentiable, i.e the Hessian matrix exists at each point in -$D_f$. Then $f$ is convex if and only if $D_f$ is a convex set and its -Hessian is positive semi-definite for all $x\in D_f$. For a -single-variable function this reduces to $f''(x) \geq -0$. Geometrically this means that $f$ has nonnegative curvature -everywhere. -!eblock - -This condition is particularly useful since it gives us an procedure for determining if the function under consideration is convex, apart from using the definition. - -!split -===== More on convex functions ===== - -The next result is of great importance to us and the reason why we are -going on about convex functions. In machine learning we frequently -have to minimize a loss/cost function in order to find the best -parameters for the model we are considering. - -Ideally we want the -global minimum (for high-dimensional models it is hard to know -if we have local or global minimum). However, if the cost/loss function -is convex the following result provides invaluable information: - -!bblock Any minimum is global for convex functions -Consider the problem of finding $x \in \mathbb{R}^n$ such that $f(x)$ -is minimal, where $f$ is convex and differentiable. Then, any point -$x^*$ that satisfies $\nabla f(x^*) = 0$ is a global minimum. -!eblock - -This result means that if we know that the cost/loss function is convex and we are able to find a minimum, we are guaranteed that it is a global minimum. - -!split -===== Some simple problems ===== - -o Show that $f(x)=x^2$ is convex for $x \in \mathbb{R}$ using the definition of convexity. Hint: If you re-write the definition, $f$ is convex if the following holds for all $x,y \in D_f$ and any $\lambda \in [0,1] $ $\lambda f(x) + (1-\lambda)f(y) - f(\lambda x + (1-\lambda) y ) \geq 0. $ - -o Using the second order condition show that the following functions are convex on the specified domain. - * $f(x) = e^x$ is convex for $x \in \mathbb{R}$. - * $g(x) = -\ln(x)$ is convex for $x \in (0,\infty)$. -o Let $f(x) = x^2$ and $g(x) = e^x$. Show that $f(g(x))$ and $g(f(x))$ is convex for $x \in \mathbb{R}$. Also show that if $f(x)$ is any convex function than $h(x) = e^{f(x)}$ is convex. - -o A norm is any function that satisfy the following properties - * $f(\alpha x) = |\alpha| f(x)$ for all $\alpha \in \mathbb{R}$. - * $f(x+y) \leq f(x) + f(y)$ - * $f(x) \leq 0$ for all $x \in \mathbb{R}^n$ with equality if and only if $x = 0$ - -Using the definition of convexity, try to show that a function satisfying the properties above is convex (the third condition is not needed to show this). - -!split -===== Revisiting our first homework ===== - -We will use linear regression as a case study for the gradient descent -methods. Linear regression is a great test case for the gradient -descent methods discussed in the lectures since it has several -desirable properties such as: - -o An analytical solution (recall homework set 1). -o The gradient can be computed analytically. -o The cost function is convex which guarantees that gradient descent converges for small enough learning rates - -We revisit the example from homework set 1 where we had -!bt -\[ -y_i = 5x_i^2 + 0.1\xi_i, \ i=1,\cdots,100 -\] -!et -with $x_i \in [0,1] $ chosen randomly with a uniform distribution. Additionally $\xi_i$ represents stochastic noise chosen according to a normal distribution $\cal {N}(0,1)$. -The linear regression model is given by -!bt -\[ -h_\beta(x) = \hat{y} = \beta_0 + \beta_1 x, -\] -!et -such that -!bt -\[ -\hat{y}_i = \beta_0 + \beta_1 x_i. -\] -!et - -!split -===== Gradient descent example ===== - -Let $\mathbf{y} = (y_1,\cdots,y_n)^T$, $\mathbf{\hat{y}} = (\hat{y}_1,\cdots,\hat{y}_n)^T$ and $\beta = (\beta_0, \beta_1)^T$ - -t is convenient to write $\mathbf{\hat{y}} = X\beta$ where $X \in \mathbb{R}^{100 \times 2} $ is the design matrix given by -!bt -\[ -\begin{equation} -X \equiv \begin{bmatrix} -1 & x_1 \\ -\vdots & \vdots \\ -1 & x_{100} & \\ -\end{bmatrix}. -\end{equation} -\] -!et -The loss function is given by -!bt -\[ -C(\beta) = ||X\beta-\mathbf{y}||^2 = ||X\beta||^2 - 2 \mathbf{y}^T X\beta + ||\mathbf{y}||^2 = \sum_{i=1}^{100} (\beta_0 + \beta_1 x_i)^2 - 2 y_i (\beta_0 + \beta_1 x_i) + y_i^2 -\] -!et -and we want to find $\beta$ such that $C(\beta)$ is minimized. - -!split -===== The derivative of the cost/loss function ===== - -Computing $\partial C(\beta) / \partial \beta_0$ and $\partial C(\beta) / \partial \beta_1$ we can show that the gradient can be written as -!bt -\[ -\nabla_\beta C(\beta) = (\partial C(\beta) / \partial \beta_0, \partial C(\beta) / \partial \beta_1)^T = 2\begin{bmatrix} \sum_{i=1}^{100} \left(\beta_0+\beta_1x_i-y_i\right) \\ -\sum_{i=1}^{100}\left( x_i (\beta_0+\beta_1x_i)-y_ix_i\right) \\ -\end{bmatrix} = 2X^T(X\beta - \mathbf{y}), -\] -!et -where $X$ is the design matrix defined above. - -!split -===== The Hessian matrix ===== -The Hessian matrix of $C(\beta)$ is given by -!bt -\[ -\hat{H} \equiv \begin{bmatrix} -\frac{\partial^2 C(\beta)}{\partial \beta_0^2} & \frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} \\ -\frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} & \frac{\partial^2 C(\beta)}{\partial \beta_1^2} & \\ -\end{bmatrix} = 2X^T X. -\] -!et -This result implies that $C(\beta)$ is a convex function since the matrix $X^T X$ always is positive semi-definite. - -!split -===== Simple program ===== - -We can now write a program that minimizes $C(\beta)$ using the gradient descent method with a constant learning rate $\gamma$ according to -!bt -\[ -\beta_{k+1} = \beta_k - \gamma \nabla_\beta C(\beta_k), \ k=0,1,\cdots -\] -!et - -We can use the expression we computed for the gradient and let use a -$\beta_0$ be chosen randomly and let $\gamma = 0.001$. Stop iterating -when $||\nabla_\beta C(\beta_k) || < \epsilon = 10^{-8}$. - -And finally we can compare our solution for $\beta$ with the analytic result given by -$\beta= (X^TX)^{-1} X^T \mathbf{y}$. -!bc pycod -import numpy as np - -""" -The following setup is just a suggestion, feel free to write it the way you like. -""" - -#Setup problem described in the exercise -N = 100 #Nr of datapoints -M = 2 #Nr of features -x = np.random.rand(N) #Uniformly generated x-values in [0,1] -y = 5*x**2 + 0.1*np.random.randn(N) -X = np.c_[np.ones(N),x] #Construct design matrix - -#Compute beta according to normal equations to compare with GD solution -Xt_X_inv = np.linalg.inv(np.dot(X.T,X)) -Xt_y = np.dot(X.transpose(),y) -beta_NE = np.dot(Xt_X_inv,Xt_y) -print(beta_NE) -!ec !split ===== Gradient descent and Ridge ===== @@ -432,7 +434,7 @@ the shortcomings of the Gradient descent method discussed above. The underlying idea of SGD comes from the observation that the cost function, which we want to minimize, can almost always be written as a -sum over $n$ datapoints $\{\mathbf{x}_i\}_{i=1}^n$, +sum over $n$ data points $\{\mathbf{x}_i\}_{i=1}^n$, !bt \[ C(\mathbf{\beta}) = \sum_{i=1}^n c_i(\mathbf{x}_i, @@ -454,27 +456,27 @@ computed as a sum over $i$-gradients Stochasticity/randomness is introduced by only taking the gradient on a subset of the data called minibatches. If there are $n$ -datapoints and the size of each minibatch is $M$, there will be $n/M$ +data points and the size of each minibatch is $M$, there will be $n/M$ minibatches. We denote these minibatches by $B_k$ where $k=1,\cdots,n/M$. !split ===== SGD example ===== -As an example, suppose we have $10$ datapoints $( \mathbf{x}_1, -\cdots, \mathbf{x}_{10} )$ and we choose to have $M=5$ minibathces, -then each minibatch contains two datapoints. In particular we have +As an example, suppose we have $10$ data points $(\mathbf{x}_1,\cdots, \mathbf{x}_{10})$ +and we choose to have $M=5$ minibathces, +then each minibatch contains two data points. In particular we have $B_1 = (\mathbf{x}_1,\mathbf{x}_2), \cdots, B_5 = (\mathbf{x}_9,\mathbf{x}_{10})$. Note that if you choose $M=1$ you -have only a single batch with all datapoints and on the other extreme, +have only a single batch with all data points and on the other extreme, you may choose $M=n$ resulting in a minibatch for each datapoint, i.e $B_k = \mathbf{x}_k$. The idea is now to approximate the gradient by replacing the sum over -all datapoints with a sum over the datapoints in one the minibatches +all data points with a sum over the data points in one the minibatches picked at random in each gradient descent step !bt \[ -\nabla_\beta +\nabla_{\beta} C(\mathbf{\beta}) = \sum_{i=1}^n \nabla_\beta c_i(\mathbf{x}_i, \mathbf{\beta}) \rightarrow \sum_{i \in B_k}^n \nabla_\beta c_i(\mathbf{x}_i, \mathbf{\beta}). @@ -522,8 +524,8 @@ Taking the gradient only on a subset of the data has two important benefits. First, it introduces randomness which decreases the chance that our opmization scheme gets stuck in a local minima. Second, if the size of the minibatches are small relative to the number of -datapoints ($M < n$), the computation of the gradient is much -cheaper since we sum over the datapoints in the k-th minibatch and not +datapoints ($M < n$), the computation of the gradient is much +cheaper since we sum over the datapoints in the $k-th$ minibatch and not all $n$ datapoints. !split @@ -547,7 +549,7 @@ Another approach is to let the step length $\gamma_j$ depend on the number of epochs in such a way that it becomes very small after a reasonable time such that we do not move at all. -As an example, let $e = 0,1,2,3,\cdots$ denote the current epoch and let $t_0, t_1 > 0$ be two fixed numbers. Furthermore, let $t = e \cdot m + i$ where $m$ is the number of minibatches and $i=0,\cdots,m-1$. Then the function $$\gamma_j(t; t_0, t_1) = \frac{t_0}{t+t_1} $$ goes to zero as the number of epochs gets large. I.e. we start with a step length $\gamma_j (0; t_0, t_1) = t_0/t_1$ which decays in *time* $t$. +As an example, let $e = 0,1,2,3,\cdots$ denote the current epoch and let $t_0, t_1 > 0$ be two fixed numbers. Furthermore, let $t = e \cdot m + i$ where $m$ is the number of minibatches and $i=0,\cdots,m-1$. Then the function $$\gamma_j(t; t_0, t_1) = \frac{t_0}{t+t_1} $$ goes to zero as the number of epochs gets large. I.e. we start with a step length $\gamma_j (0; t_0, t_1) = t_0/t_1$ which decays in *time* $t$. In this way we can fix the number of epochs, compute $\beta$ and evaluate the cost function at the end. Repeating the computation will