diff --git a/doc/pub/week40/html/._week40-bs000.html b/doc/pub/week40/html/._week40-bs000.html index 76cb74266..dd9e9bdc5 100644 --- a/doc/pub/week40/html/._week40-bs000.html +++ b/doc/pub/week40/html/._week40-bs000.html @@ -63,6 +63,7 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d 2, None, 'program-for-stochastic-gradient'), + ('Replace or not', 2, None, 'replace-or-not'), ('Momentum based GD', 2, None, 'momentum-based-gd'), ('More on momentum based approaches', 2, @@ -249,62 +250,63 @@ MathJax.Hub.Config({
-
The stochastic gradient descent (SGD) is almost always used with a -momentum or inertia term that serves as a memory of the direction we -are moving in parameter space. This is typically implemented as -follows +
In the above code, we have use replacement in setting up the +mini-batches. The discussion +here may be +useful. More material will be added later.
-$$ -\begin{align} -\mathbf{v}_{t}&=\gamma \mathbf{v}_{t-1}+\eta_{t}\nabla_\theta E(\boldsymbol{\theta}_t) \nonumber \\ -\boldsymbol{\theta}_{t+1}&= \boldsymbol{\theta}_t -\mathbf{v}_{t}, -\tag{1} -\end{align} -$$ - -where we have introduced a momentum parameter \( \gamma \), with -\( 0\le\gamma\le 1 \), and for brevity we dropped the explicit notation to -indicate the gradient is to be taken over a different mini-batch at -each step. We call this algorithm gradient descent with momentum -(GDM). From these equations, it is clear that \( \mathbf{v}_t \) is a -running average of recently encountered gradients and -\( (1-\gamma)^{-1} \) sets the characteristic time scale for the memory -used in the averaging procedure. Consistent with this, when -\( \gamma=0 \), this just reduces down to ordinary SGD as discussed -earlier. An equivalent way of writing the updates is -
- -$$ -\Delta \boldsymbol{\theta}_{t+1} = \gamma \Delta \boldsymbol{\theta}_t -\ \eta_{t}\nabla_\theta E(\boldsymbol{\theta}_t), -$$ - -where we have defined \( \Delta \boldsymbol{\theta}_{t}= \boldsymbol{\theta}_t-\boldsymbol{\theta}_{t-1} \).
-diff --git a/doc/pub/week40/html/._week40-bs014.html b/doc/pub/week40/html/._week40-bs014.html index 019aefc34..2889731d2 100644 --- a/doc/pub/week40/html/._week40-bs014.html +++ b/doc/pub/week40/html/._week40-bs014.html @@ -63,6 +63,7 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d 2, None, 'program-for-stochastic-gradient'), + ('Replace or not', 2, None, 'replace-or-not'), ('Momentum based GD', 2, None, 'momentum-based-gd'), ('More on momentum based approaches', 2, @@ -249,62 +250,63 @@ MathJax.Hub.Config({
-
Let us try to get more intuition from these equations. It is helpful -to consider a simple physical analogy with a particle of mass \( m \) -moving in a viscous medium with drag coefficient \( \mu \) and potential -\( E(\mathbf{w}) \). If we denote the particle's position by \( \mathbf{w} \), -then its motion is described by +
The stochastic gradient descent (SGD) is almost always used with a +momentum or inertia term that serves as a memory of the direction we +are moving in parameter space. This is typically implemented as +follows
$$ -m {d^2 \mathbf{w} \over dt^2} + \mu {d \mathbf{w} \over dt }= -\nabla_w E(\mathbf{w}). +\begin{align} +\mathbf{v}_{t}&=\gamma \mathbf{v}_{t-1}+\eta_{t}\nabla_\theta E(\boldsymbol{\theta}_t) \nonumber \\ +\boldsymbol{\theta}_{t+1}&= \boldsymbol{\theta}_t -\mathbf{v}_{t}, +\tag{1} +\end{align} $$ -We can discretize this equation in the usual way to get
+where we have introduced a momentum parameter \( \gamma \), with +\( 0\le\gamma\le 1 \), and for brevity we dropped the explicit notation to +indicate the gradient is to be taken over a different mini-batch at +each step. We call this algorithm gradient descent with momentum +(GDM). From these equations, it is clear that \( \mathbf{v}_t \) is a +running average of recently encountered gradients and +\( (1-\gamma)^{-1} \) sets the characteristic time scale for the memory +used in the averaging procedure. Consistent with this, when +\( \gamma=0 \), this just reduces down to ordinary SGD as discussed +earlier. An equivalent way of writing the updates is +
$$ -m { \mathbf{w}_{t+\Delta t}-2 \mathbf{w}_{t} +\mathbf{w}_{t-\Delta t} \over (\Delta t)^2}+\mu {\mathbf{w}_{t+\Delta t}- \mathbf{w}_{t} \over \Delta t} = -\nabla_w E(\mathbf{w}). -$$ - -Rearranging this equation, we can rewrite this as
- -$$ -\Delta \mathbf{w}_{t +\Delta t}= - { (\Delta t)^2 \over m +\mu \Delta t} \nabla_w E(\mathbf{w})+ {m \over m +\mu \Delta t} \Delta \mathbf{w}_t. +\Delta \boldsymbol{\theta}_{t+1} = \gamma \Delta \boldsymbol{\theta}_t -\ \eta_{t}\nabla_\theta E(\boldsymbol{\theta}_t), $$ +where we have defined \( \Delta \boldsymbol{\theta}_{t}= \boldsymbol{\theta}_t-\boldsymbol{\theta}_{t-1} \).
@@ -367,7 +377,7 @@ $$
-
Notice that this equation is identical to previous one if we identify -the position of the particle, \( \mathbf{w} \), with the parameters -\( \boldsymbol{\theta} \). This allows us to identify the momentum -parameter and learning rate with the mass of the particle and the -viscous drag as: +
Let us try to get more intuition from these equations. It is helpful +to consider a simple physical analogy with a particle of mass \( m \) +moving in a viscous medium with drag coefficient \( \mu \) and potential +\( E(\mathbf{w}) \). If we denote the particle's position by \( \mathbf{w} \), +then its motion is described by
$$ -\gamma= {m \over m +\mu \Delta t }, \qquad \eta = {(\Delta t)^2 \over m +\mu \Delta t}. +m {d^2 \mathbf{w} \over dt^2} + \mu {d \mathbf{w} \over dt }= -\nabla_w E(\mathbf{w}). $$ -Thus, as the name suggests, the momentum parameter is proportional to -the mass of the particle and effectively provides inertia. -Furthermore, in the large viscosity/small learning rate limit, our -memory time scales as \( (1-\gamma)^{-1} \approx m/(\mu \Delta t) \). -
- -Why is momentum useful? SGD momentum helps the gradient descent -algorithm gain speed in directions with persistent but small gradients -even in the presence of stochasticity, while suppressing oscillations -in high-curvature directions. This becomes especially important in -situations where the landscape is shallow and flat in some directions -and narrow and steep in others. It has been argued that first-order -methods (with appropriate initial conditions) can perform comparable -to more expensive second order methods, especially in the context of -complex deep learning models. -
- -These beneficial properties of momentum can sometimes become even more -pronounced by using a slight modification of the classical momentum -algorithm called Nesterov Accelerated Gradient (NAG). -
- -In the NAG algorithm, rather than calculating the gradient at the -current parameters, \( \nabla_\theta E(\boldsymbol{\theta}_t) \), one -calculates the gradient at the expected value of the parameters given -our current momentum, \( \nabla_\theta E(\boldsymbol{\theta}_t +\gamma -\mathbf{v}_{t-1}) \). This yields the NAG update rule -
+We can discretize this equation in the usual way to get
$$ -\begin{align} -\mathbf{v}_{t}&=\gamma \mathbf{v}_{t-1}+\eta_{t}\nabla_\theta E(\boldsymbol{\theta}_t +\gamma \mathbf{v}_{t-1}) \nonumber \\ -\boldsymbol{\theta}_{t+1}&= \boldsymbol{\theta}_t -\mathbf{v}_{t}. -\tag{2} -\end{align} +m { \mathbf{w}_{t+\Delta t}-2 \mathbf{w}_{t} +\mathbf{w}_{t-\Delta t} \over (\Delta t)^2}+\mu {\mathbf{w}_{t+\Delta t}- \mathbf{w}_{t} \over \Delta t} = -\nabla_w E(\mathbf{w}). +$$ + +Rearranging this equation, we can rewrite this as
+ +$$ +\Delta \mathbf{w}_{t +\Delta t}= - { (\Delta t)^2 \over m +\mu \Delta t} \nabla_w E(\mathbf{w})+ {m \over m +\mu \Delta t} \Delta \mathbf{w}_t. $$ -One of the major advantages of NAG is that it allows for the use of a larger learning rate than GDM for the same choice of \( \gamma \).
@@ -393,7 +369,7 @@ $$
-
In stochastic gradient descent, with and without momentum, we still -have to specify a schedule for tuning the learning rates \( \eta_t \) -as a function of time. As discussed in the context of Newton's -method, this presents a number of dilemmas. The learning rate is -limited by the steepest direction which can change depending on the -current position in the landscape. To circumvent this problem, ideally -our algorithm would keep track of curvature and take large steps in -shallow, flat directions and small steps in steep, narrow directions. -Second-order methods accomplish this by calculating or approximating -the Hessian and normalizing the learning rate by the -curvature. However, this is very computationally expensive for -extremely large models. Ideally, we would like to be able to -adaptively change the step size to match the landscape without paying -the steep computational price of calculating or approximating -Hessians. +
Notice that this equation is identical to previous one if we identify +the position of the particle, \( \mathbf{w} \), with the parameters +\( \boldsymbol{\theta} \). This allows us to identify the momentum +parameter and learning rate with the mass of the particle and the +viscous drag as:
-Recently, a number of methods have been introduced that accomplish -this by tracking not only the gradient, but also the second moment of -the gradient. These methods include AdaGrad, AdaDelta, Root Mean Squared Propagation (RMS-Prop), and -ADAM. +$$ +\gamma= {m \over m +\mu \Delta t }, \qquad \eta = {(\Delta t)^2 \over m +\mu \Delta t}. +$$ + +
Thus, as the name suggests, the momentum parameter is proportional to +the mass of the particle and effectively provides inertia. +Furthermore, in the large viscosity/small learning rate limit, our +memory time scales as \( (1-\gamma)^{-1} \approx m/(\mu \Delta t) \).
+Why is momentum useful? SGD momentum helps the gradient descent +algorithm gain speed in directions with persistent but small gradients +even in the presence of stochasticity, while suppressing oscillations +in high-curvature directions. This becomes especially important in +situations where the landscape is shallow and flat in some directions +and narrow and steep in others. It has been argued that first-order +methods (with appropriate initial conditions) can perform comparable +to more expensive second order methods, especially in the context of +complex deep learning models. +
+ +These beneficial properties of momentum can sometimes become even more +pronounced by using a slight modification of the classical momentum +algorithm called Nesterov Accelerated Gradient (NAG). +
+ +In the NAG algorithm, rather than calculating the gradient at the +current parameters, \( \nabla_\theta E(\boldsymbol{\theta}_t) \), one +calculates the gradient at the expected value of the parameters given +our current momentum, \( \nabla_\theta E(\boldsymbol{\theta}_t +\gamma +\mathbf{v}_{t-1}) \). This yields the NAG update rule +
+ +$$ +\begin{align} +\mathbf{v}_{t}&=\gamma \mathbf{v}_{t-1}+\eta_{t}\nabla_\theta E(\boldsymbol{\theta}_t +\gamma \mathbf{v}_{t-1}) \nonumber \\ +\boldsymbol{\theta}_{t+1}&= \boldsymbol{\theta}_t -\mathbf{v}_{t}. +\tag{2} +\end{align} +$$ + +One of the major advantages of NAG is that it allows for the use of a larger learning rate than GDM for the same choice of \( \gamma \).
+diff --git a/doc/pub/week40/html/._week40-bs017.html b/doc/pub/week40/html/._week40-bs017.html index f1d908fc6..b6c61f4fd 100644 --- a/doc/pub/week40/html/._week40-bs017.html +++ b/doc/pub/week40/html/._week40-bs017.html @@ -63,6 +63,7 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d 2, None, 'program-for-stochastic-gradient'), + ('Replace or not', 2, None, 'replace-or-not'), ('Momentum based GD', 2, None, 'momentum-based-gd'), ('More on momentum based approaches', 2, @@ -249,62 +250,63 @@ MathJax.Hub.Config({
-
In RMS prop, in addition to keeping a running average of the first -moment of the gradient, we also keep track of the second moment -denoted by \( \mathbf{s}_t=\mathbb{E}[\mathbf{g}_t^2] \). The update rule -for RMS prop is given by +
In stochastic gradient descent, with and without momentum, we still +have to specify a schedule for tuning the learning rates \( \eta_t \) +as a function of time. As discussed in the context of Newton's +method, this presents a number of dilemmas. The learning rate is +limited by the steepest direction which can change depending on the +current position in the landscape. To circumvent this problem, ideally +our algorithm would keep track of curvature and take large steps in +shallow, flat directions and small steps in steep, narrow directions. +Second-order methods accomplish this by calculating or approximating +the Hessian and normalizing the learning rate by the +curvature. However, this is very computationally expensive for +extremely large models. Ideally, we would like to be able to +adaptively change the step size to match the landscape without paying +the steep computational price of calculating or approximating +Hessians.
-$$ -\begin{align} -\mathbf{g}_t &= \nabla_\theta E(\boldsymbol{\theta}) -\tag{3}\\ -\mathbf{s}_t &=\beta \mathbf{s}_{t-1} +(1-\beta)\mathbf{g}_t^2 \nonumber \\ -\boldsymbol{\theta}_{t+1}&=&\boldsymbol{\theta}_t - \eta_t { \mathbf{g}_t \over \sqrt{\mathbf{s}_t +\epsilon}}, \nonumber -\end{align} -$$ - -where \( \beta \) controls the averaging time of the second moment and is -typically taken to be about \( \beta=0.9 \), \( \eta_t \) is a learning rate -typically chosen to be \( 10^{-3} \), and \( \epsilon\sim 10^{-8} \) is a -small regularization constant to prevent divergences. Multiplication -and division by vectors is understood as an element-wise operation. It -is clear from this formula that the learning rate is reduced in -directions where the norm of the gradient is consistently large. This -greatly speeds up the convergence by allowing us to use a larger -learning rate for flat directions. +
Recently, a number of methods have been introduced that accomplish +this by tracking not only the gradient, but also the second moment of +the gradient. These methods include AdaGrad, AdaDelta, Root Mean Squared Propagation (RMS-Prop), and +ADAM.
@@ -369,7 +368,7 @@ learning rate for flat directions.
-
A related algorithm is the ADAM optimizer. In ADAM, we keep a running -average of both the first and second moment of the gradient and use -this information to adaptively change the learning rate for different -parameters. In addition to keeping a running average of the first and -second moments of the gradient -(i.e. \( \mathbf{m}_t=\mathbb{E}[\mathbf{g}_t] \) and -\( \mathbf{s}_t=\mathbb{E}[\mathbf{g}^2_t] \), respectively), ADAM -performs an additional bias correction to account for the fact that we -are estimating the first two moments of the gradient using a running -average (denoted by the hats in the update rule below). The update -rule for ADAM is given by (where multiplication and division are once -again understood to be element-wise operations below) +
In RMS prop, in addition to keeping a running average of the first +moment of the gradient, we also keep track of the second moment +denoted by \( \mathbf{s}_t=\mathbb{E}[\mathbf{g}_t^2] \). The update rule +for RMS prop is given by
$$ \begin{align} \mathbf{g}_t &= \nabla_\theta E(\boldsymbol{\theta}) -\tag{4}\\ -\mathbf{m}_t &= \beta_1 \mathbf{m}_{t-1} + (1-\beta_1) \mathbf{g}_t \nonumber \\ -\mathbf{s}_t &=\beta_2 \mathbf{s}_{t-1} +(1-\beta_2)\mathbf{g}_t^2 \nonumber \\ -\boldsymbol{\mathbf{m}}_t&={\mathbf{m}_t \over 1-\beta_1^t} \nonumber \\ -\boldsymbol{\mathbf{s}}_t &={\mathbf{s}_t \over1-\beta_2^t} \nonumber \\ -\boldsymbol{\theta}_{t+1}&=\boldsymbol{\theta}_t - \eta_t { \boldsymbol{\mathbf{m}}_t \over \sqrt{\boldsymbol{\mathbf{s}}_t} +\epsilon}, \nonumber \\ -\tag{5} +\tag{3}\\ +\mathbf{s}_t &=\beta \mathbf{s}_{t-1} +(1-\beta)\mathbf{g}_t^2 \nonumber \\ +\boldsymbol{\theta}_{t+1}&=&\boldsymbol{\theta}_t - \eta_t { \mathbf{g}_t \over \sqrt{\mathbf{s}_t +\epsilon}}, \nonumber \end{align} $$ -where \( \beta_1 \) and \( \beta_2 \) set the memory lifetime of the first and -second moment and are typically taken to be \( 0.9 \) and \( 0.99 \) -respectively, and \( \eta \) and \( \epsilon \) are identical to RMSprop. +
where \( \beta \) controls the averaging time of the second moment and is +typically taken to be about \( \beta=0.9 \), \( \eta_t \) is a learning rate +typically chosen to be \( 10^{-3} \), and \( \epsilon\sim 10^{-8} \) is a +small regularization constant to prevent divergences. Multiplication +and division by vectors is understood as an element-wise operation. It +is clear from this formula that the learning rate is reduced in +directions where the norm of the gradient is consistently large. This +greatly speeds up the convergence by allowing us to use a larger +learning rate for flat directions.
-Like in RMSprop, the effective step size of a parameter depends on the -magnitude of its gradient squared. To understand this better, let us -rewrite this expression in terms of the variance -\( \boldsymbol{\sigma}_t^2 = \boldsymbol{\mathbf{s}}_t - -(\boldsymbol{\mathbf{m}}_t)^2 \). Consider a single parameter \( \theta_t \). The -update rule for this parameter is given by -
- -$$ -\Delta \theta_{t+1}= -\eta_t { \boldsymbol{m}_t \over \sqrt{\sigma_t^2 + m_t^2 }+\epsilon}. -$$ - -diff --git a/doc/pub/week40/html/._week40-bs019.html b/doc/pub/week40/html/._week40-bs019.html index dece297a4..4b4a8d444 100644 --- a/doc/pub/week40/html/._week40-bs019.html +++ b/doc/pub/week40/html/._week40-bs019.html @@ -63,6 +63,7 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d 2, None, 'program-for-stochastic-gradient'), + ('Replace or not', 2, None, 'replace-or-not'), ('Momentum based GD', 2, None, 'momentum-based-gd'), ('More on momentum based approaches', 2, @@ -249,62 +250,63 @@ MathJax.Hub.Config({
-
A related algorithm is the ADAM optimizer. In ADAM, we keep a running +average of both the first and second moment of the gradient and use +this information to adaptively change the learning rate for different +parameters. In addition to keeping a running average of the first and +second moments of the gradient +(i.e. \( \mathbf{m}_t=\mathbb{E}[\mathbf{g}_t] \) and +\( \mathbf{s}_t=\mathbb{E}[\mathbf{g}^2_t] \), respectively), ADAM +performs an additional bias correction to account for the fact that we +are estimating the first two moments of the gradient using a running +average (denoted by the hats in the update rule below). The update +rule for ADAM is given by (where multiplication and division are once +again understood to be element-wise operations below) +
+ +$$ +\begin{align} +\mathbf{g}_t &= \nabla_\theta E(\boldsymbol{\theta}) +\tag{4}\\ +\mathbf{m}_t &= \beta_1 \mathbf{m}_{t-1} + (1-\beta_1) \mathbf{g}_t \nonumber \\ +\mathbf{s}_t &=\beta_2 \mathbf{s}_{t-1} +(1-\beta_2)\mathbf{g}_t^2 \nonumber \\ +\boldsymbol{\mathbf{m}}_t&={\mathbf{m}_t \over 1-\beta_1^t} \nonumber \\ +\boldsymbol{\mathbf{s}}_t &={\mathbf{s}_t \over1-\beta_2^t} \nonumber \\ +\boldsymbol{\theta}_{t+1}&=\boldsymbol{\theta}_t - \eta_t { \boldsymbol{\mathbf{m}}_t \over \sqrt{\boldsymbol{\mathbf{s}}_t} +\epsilon}, \nonumber \\ +\tag{5} +\end{align} +$$ + +where \( \beta_1 \) and \( \beta_2 \) set the memory lifetime of the first and +second moment and are typically taken to be \( 0.9 \) and \( 0.99 \) +respectively, and \( \eta \) and \( \epsilon \) are identical to RMSprop. +
+ +Like in RMSprop, the effective step size of a parameter depends on the +magnitude of its gradient squared. To understand this better, let us +rewrite this expression in terms of the variance +\( \boldsymbol{\sigma}_t^2 = \boldsymbol{\mathbf{s}}_t - +(\boldsymbol{\mathbf{m}}_t)^2 \). Consider a single parameter \( \theta_t \). The +update rule for this parameter is given by +
+ +$$ +\Delta \theta_{t+1}= -\eta_t { \boldsymbol{m}_t \over \sqrt{\sigma_t^2 + m_t^2 }+\epsilon}. +$$ -Geron's text, see chapter 11, has several interesting discussions.
@@ -351,7 +390,7 @@ MathJax.Hub.Config({
-
Automatic differentiation (AD), -also called algorithmic -differentiation or computational differentiation,is a set of -techniques to numerically evaluate the derivative of a function -specified by a computer program. AD exploits the fact that every -computer program, no matter how complicated, executes a sequence of -elementary arithmetic operations (addition, subtraction, -multiplication, division, etc.) and elementary functions (exp, log, -sin, cos, etc.). By applying the chain rule repeatedly to these -operations, derivatives of arbitrary order can be computed -automatically, accurately to working precision, and using at most a -small constant factor more arithmetic operations than the original -program. -
- -Automatic differentiation is neither:
+Symbolic differentiation can lead to inefficient code and faces the -difficulty of converting a computer program into a single expression, -while numerical differentiation can introduce round-off errors in the -discretization process and cancellation -
- -Python has tools for so-called automatic differentiation. -Consider the following example -
-$$ -f(x) = \sin\left(2\pi x + x^2\right) -$$ - -which has the following derivative
-$$ -f'(x) = \cos\left(2\pi x + x^2\right)\left(2\pi + 2x\right) -$$ - -Using autograd we have
- - - -import autograd.numpy as np
-
-# To do elementwise differentiation:
-from autograd import elementwise_grad as egrad
-
-# To plot:
-import matplotlib.pyplot as plt
-
-
-def f(x):
- return np.sin(2*np.pi*x + x**2)
-
-def f_grad_analytic(x):
- return np.cos(2*np.pi*x + x**2)*(2*np.pi + 2*x)
-
-# Do the comparison:
-x = np.linspace(0,1,1000)
-
-f_grad = egrad(f)
-
-computed = f_grad(x)
-analytic = f_grad_analytic(x)
-
-plt.title('Derivative computed from Autograd compared with the analytical derivative')
-plt.plot(x,computed,label='autograd')
-plt.plot(x,analytic,label='analytic')
-
-plt.xlabel('x')
-plt.ylabel('y')
-plt.legend()
-
-plt.show()
-
-print("The max absolute difference is: %g"%(np.max(np.abs(computed - analytic))))
-
-Geron's text, see chapter 11, has several interesting discussions.
@@ -441,7 +353,7 @@ plt.show()
- -
Here we -experiment with what kind of functions Autograd is capable -of finding the gradient of. The following Python functions are just -meant to illustrate what Autograd can do, but please feel free to -experiment with other, possibly more complicated, functions as well. +
Automatic differentiation (AD), +also called algorithmic +differentiation or computational differentiation,is a set of +techniques to numerically evaluate the derivative of a function +specified by a computer program. AD exploits the fact that every +computer program, no matter how complicated, executes a sequence of +elementary arithmetic operations (addition, subtraction, +multiplication, division, etc.) and elementary functions (exp, log, +sin, cos, etc.). By applying the chain rule repeatedly to these +operations, derivatives of arbitrary order can be computed +automatically, accurately to working precision, and using at most a +small constant factor more arithmetic operations than the original +program.
+Automatic differentiation is neither:
+ +Symbolic differentiation can lead to inefficient code and faces the +difficulty of converting a computer program into a single expression, +while numerical differentiation can introduce round-off errors in the +discretization process and cancellation +
+ +Python has tools for so-called automatic differentiation. +Consider the following example +
+$$ +f(x) = \sin\left(2\pi x + x^2\right) +$$ + +which has the following derivative
+$$ +f'(x) = \cos\left(2\pi x + x^2\right)\left(2\pi + 2x\right) +$$ + +Using autograd we have
+import autograd.numpy as np
-from autograd import grad
-def f1(x):
- return x**3 + 1
+# To do elementwise differentiation:
+from autograd import elementwise_grad as egrad
-f1_grad = grad(f1)
+# To plot:
+import matplotlib.pyplot as plt
-# Remember to send in float as argument to the computed gradient from Autograd!
-a = 1.0
-# See the evaluated gradient at a using autograd:
-print("The gradient of f1 evaluated at a = %g using autograd is: %g"%(a,f1_grad(a)))
+def f(x):
+ return np.sin(2*np.pi*x + x**2)
-# Compare with the analytical derivative, that is f1'(x) = 3*x**2
-grad_analytical = 3*a**2
-print("The gradient of f1 evaluated at a = %g by finding the analytic expression is: %g"%(a,grad_analytical))
+def f_grad_analytic(x):
+ return np.cos(2*np.pi*x + x**2)*(2*np.pi + 2*x)
+
+# Do the comparison:
+x = np.linspace(0,1,1000)
+
+f_grad = egrad(f)
+
+computed = f_grad(x)
+analytic = f_grad_analytic(x)
+
+plt.title('Derivative computed from Autograd compared with the analytical derivative')
+plt.plot(x,computed,label='autograd')
+plt.plot(x,analytic,label='analytic')
+
+plt.xlabel('x')
+plt.ylabel('y')
+plt.legend()
+
+plt.show()
+
+print("The max absolute difference is: %g"%(np.max(np.abs(computed - analytic))))
- -
To differentiate with respect to two (or more) arguments of a Python -function, Autograd need to know at which variable the function if -being differentiated with respect to. +
Here we +experiment with what kind of functions Autograd is capable +of finding the gradient of. The following Python functions are just +meant to illustrate what Autograd can do, but please feel free to +experiment with other, possibly more complicated, functions as well.
@@ -332,37 +336,21 @@ being differentiated with respect to.import autograd.numpy as np
from autograd import grad
-def f2(x1,x2):
- return 3*x1**3 + x2*(x1 - 5) + 1
-# By sending the argument 0, Autograd will compute the derivative w.r.t the first variable, in this case x1
-f2_grad_x1 = grad(f2,0)
+def f1(x):
+ return x**3 + 1
-# ... and differentiate w.r.t x2 by sending 1 as an additional arugment to grad
-f2_grad_x2 = grad(f2,1)
+f1_grad = grad(f1)
-x1 = 1.0
-x2 = 3.0
+# Remember to send in float as argument to the computed gradient from Autograd!
+a = 1.0
-print("Evaluating at x1 = %g, x2 = %g"%(x1,x2))
-print("-"*30)
+# See the evaluated gradient at a using autograd:
+print("The gradient of f1 evaluated at a = %g using autograd is: %g"%(a,f1_grad(a)))
-# Compare with the analytical derivatives:
-
-# Derivative of f2 w.r.t x1 is: 9*x1**2 + x2:
-f2_grad_x1_analytical = 9*x1**2 + x2
-
-# Derivative of f2 w.r.t x2 is: x1 - 5:
-f2_grad_x2_analytical = x1 - 5
-
-# See the evaluated derivations:
-print("The derivative of f2 w.r.t x1: %g"%( f2_grad_x1(x1,x2) ))
-print("The analytical derivative of f2 w.r.t x1: %g"%( f2_grad_x1(x1,x2) ))
-
-print()
-
-print("The derivative of f2 w.r.t x2: %g"%( f2_grad_x2(x1,x2) ))
-print("The analytical derivative of f2 w.r.t x2: %g"%( f2_grad_x2(x1,x2) ))
+# Compare with the analytical derivative, that is f1'(x) = 3*x**2
+grad_analytical = 3*a**2
+print("The gradient of f1 evaluated at a = %g by finding the analytic expression is: %g"%(a,grad_analytical))
-
To differentiate with respect to two (or more) arguments of a Python +function, Autograd need to know at which variable the function if +being differentiated with respect to. +
@@ -327,21 +334,37 @@ MathJax.Hub.Config({import autograd.numpy as np
from autograd import grad
-def f3(x): # Assumes x is an array of length 5 or higher
- return 2*x[0] + 3*x[1] + 5*x[2] + 7*x[3] + 11*x[4]**2
+def f2(x1,x2):
+ return 3*x1**3 + x2*(x1 - 5) + 1
-f3_grad = grad(f3)
+# By sending the argument 0, Autograd will compute the derivative w.r.t the first variable, in this case x1
+f2_grad_x1 = grad(f2,0)
-x = np.linspace(0,4,5)
+# ... and differentiate w.r.t x2 by sending 1 as an additional arugment to grad
+f2_grad_x2 = grad(f2,1)
-# Print the computed gradient:
-print("The computed gradient of f3 is: ", f3_grad(x))
+x1 = 1.0
+x2 = 3.0
-# The analytical gradient is: (2, 3, 5, 7, 22*x[4])
-f3_grad_analytical = np.array([2, 3, 5, 7, 22*x[4]])
+print("Evaluating at x1 = %g, x2 = %g"%(x1,x2))
+print("-"*30)
-# Print the analytical gradient:
-print("The analytical gradient of f3 is: ", f3_grad_analytical)
+# Compare with the analytical derivatives:
+
+# Derivative of f2 w.r.t x1 is: 9*x1**2 + x2:
+f2_grad_x1_analytical = 9*x1**2 + x2
+
+# Derivative of f2 w.r.t x2 is: x1 - 5:
+f2_grad_x2_analytical = x1 - 5
+
+# See the evaluated derivations:
+print("The derivative of f2 w.r.t x1: %g"%( f2_grad_x1(x1,x2) ))
+print("The analytical derivative of f2 w.r.t x1: %g"%( f2_grad_x1(x1,x2) ))
+
+print()
+
+print("The derivative of f2 w.r.t x2: %g"%( f2_grad_x2(x1,x2) ))
+print("The analytical derivative of f2 w.r.t x2: %g"%( f2_grad_x2(x1,x2) ))
- -
import autograd.numpy as np
from autograd import grad
-def f4(x):
- return np.sqrt(1+x**2) + np.exp(x) + np.sin(2*np.pi*x)
+def f3(x): # Assumes x is an array of length 5 or higher
+ return 2*x[0] + 3*x[1] + 5*x[2] + 7*x[3] + 11*x[4]**2
-f4_grad = grad(f4)
+f3_grad = grad(f3)
-x = 2.7
+x = np.linspace(0,4,5)
-# Print the computed derivative:
-print("The computed derivative of f4 at x = %g is: %g"%(x,f4_grad(x)))
+# Print the computed gradient:
+print("The computed gradient of f3 is: ", f3_grad(x))
-# The analytical derivative is: x/sqrt(1 + x**2) + exp(x) + cos(2*pi*x)*2*pi
-f4_grad_analytical = x/np.sqrt(1 + x**2) + np.exp(x) + np.cos(2*np.pi*x)*2*np.pi
+# The analytical gradient is: (2, 3, 5, 7, 22*x[4])
+f3_grad_analytical = np.array([2, 3, 5, 7, 22*x[4]])
# Print the analytical gradient:
-print("The analytical gradient of f4 at x = %g is: %g"%(x,f4_grad_analytical))
+print("The analytical gradient of f3 is: ", f3_grad_analytical)
- -
import autograd.numpy as np
from autograd import grad
-def f5(x):
- if x >= 0:
- return x**2
- else:
- return -3*x + 1
+def f4(x):
+ return np.sqrt(1+x**2) + np.exp(x) + np.sin(2*np.pi*x)
-f5_grad = grad(f5)
+f4_grad = grad(f4)
x = 2.7
# Print the computed derivative:
-print("The computed derivative of f5 at x = %g is: %g"%(x,f5_grad(x)))
+print("The computed derivative of f4 at x = %g is: %g"%(x,f4_grad(x)))
+
+# The analytical derivative is: x/sqrt(1 + x**2) + exp(x) + cos(2*pi*x)*2*pi
+f4_grad_analytical = x/np.sqrt(1 + x**2) + np.exp(x) + np.cos(2*np.pi*x)*2*np.pi
+
+# Print the analytical gradient:
+print("The analytical gradient of f4 at x = %g is: %g"%(x,f4_grad_analytical))
-
import autograd.numpy as np
from autograd import grad
-def f6_for(x):
- val = 0
- for i in range(10):
- val = val + x**i
- return val
+def f5(x):
+ if x >= 0:
+ return x**2
+ else:
+ return -3*x + 1
-def f6_while(x):
- val = 0
- i = 0
- while i < 10:
- val = val + x**i
- i = i + 1
- return val
+f5_grad = grad(f5)
-f6_for_grad = grad(f6_for)
-f6_while_grad = grad(f6_while)
+x = 2.7
-x = 0.5
-
-# Print the computed derivaties of f6_for and f6_while
-print("The computed derivative of f6_for at x = %g is: %g"%(x,f6_for_grad(x)))
-print("The computed derivative of f6_while at x = %g is: %g"%(x,f6_while_grad(x)))
-
-import autograd.numpy as np
-from autograd import grad
-# Both of the functions are implementation of the sum: sum(x**i) for i = 0, ..., 9
-# The analytical derivative is: sum(i*x**(i-1))
-f6_grad_analytical = 0
-for i in range(10):
- f6_grad_analytical += i*x**(i-1)
-
-print("The analytical derivative of f6 at x = %g is: %g"%(x,f6_grad_analytical))
+# Print the computed derivative:
+print("The computed derivative of f5 at x = %g is: %g"%(x,f5_grad(x)))
-
import autograd.numpy as np
from autograd import grad
+def f6_for(x):
+ val = 0
+ for i in range(10):
+ val = val + x**i
+ return val
-def f7(n): # Assume that n is an integer
- if n == 1 or n == 0:
- return 1
- else:
- return n*f7(n-1)
+def f6_while(x):
+ val = 0
+ i = 0
+ while i < 10:
+ val = val + x**i
+ i = i + 1
+ return val
-f7_grad = grad(f7)
+f6_for_grad = grad(f6_for)
+f6_while_grad = grad(f6_while)
-n = 2.0
+x = 0.5
-print("The computed derivative of f7 at n = %d is: %g"%(n,f7_grad(n)))
-
-# The function f7 is an implementation of the factorial of n.
-# By using the product rule, one can find that the derivative is:
-
-f7_grad_analytical = 0
-for i in range(int(n)-1):
- tmp = 1
- for k in range(int(n)-1):
- if k != i:
- tmp *= (n - k)
- f7_grad_analytical += tmp
-
-print("The analytical derivative of f7 at n = %d is: %g"%(n,f7_grad_analytical))
+# Print the computed derivaties of f6_for and f6_while
+print("The computed derivative of f6_for at x = %g is: %g"%(x,f6_for_grad(x)))
+print("The computed derivative of f6_while at x = %g is: %g"%(x,f6_while_grad(x)))
+
+import autograd.numpy as np
+from autograd import grad
+# Both of the functions are implementation of the sum: sum(x**i) for i = 0, ..., 9
+# The analytical derivative is: sum(i*x**(i-1))
+f6_grad_analytical = 0
+for i in range(10):
+ f6_grad_analytical += i*x**(i-1)
+
+print("The analytical derivative of f6 at x = %g is: %g"%(x,f6_grad_analytical))
-
Autograd supports many features. However, there are some functions that is not supported (yet) by Autograd.
- -Assigning a value to the variable being differentiated with respect to
+import autograd.numpy as np
from autograd import grad
-def f8(x): # Assume x is an array
- x[2] = 3
- return x*2
-f8_grad = grad(f8)
+def f7(n): # Assume that n is an integer
+ if n == 1 or n == 0:
+ return 1
+ else:
+ return n*f7(n-1)
-x = 8.4
+f7_grad = grad(f7)
-print("The derivative of f8 is:",f8_grad(x))
+n = 2.0
+
+print("The computed derivative of f7 at n = %d is: %g"%(n,f7_grad(n)))
+
+# The function f7 is an implementation of the factorial of n.
+# By using the product rule, one can find that the derivative is:
+
+f7_grad_analytical = 0
+for i in range(int(n)-1):
+ tmp = 1
+ for k in range(int(n)-1):
+ if k != i:
+ tmp *= (n - k)
+ f7_grad_analytical += tmp
+
+print("The analytical derivative of f7 at n = %d is: %g"%(n,f7_grad_analytical))
Here, Autograd tells us that an 'ArrayBox' does not support item assignment. The item assignment is done when the program tries to assign x[2] to the value 3. However, Autograd has implemented the computation of the derivative such that this assignment is not possible.
+Note that if n is equal to zero or one, Autograd will give an error message. This message appears when the output is independent on input.
@@ -380,7 +395,7 @@ x = 8.4
-
Autograd supports many features. However, there are some functions that is not supported (yet) by Autograd.
+ +Assigning a value to the variable being differentiated with respect to
import autograd.numpy as np
from autograd import grad
-def f9(a): # Assume a is an array with 2 elements
- b = np.array([1.0,2.0])
- return a.dot(b)
+def f8(x): # Assume x is an array
+ x[2] = 3
+ return x*2
-f9_grad = grad(f9)
+f8_grad = grad(f8)
-x = np.array([1.0,0.0])
+x = 8.4
-print("The derivative of f9 is:",f9_grad(x))
-
-Here we are told that the 'dot' function does not belong to Autograd's -version of a Numpy array. To overcome this, an alternative syntax -which also computed the dot product can be used: -
- - - -import autograd.numpy as np
-from autograd import grad
-def f9_alternative(x): # Assume a is an array with 2 elements
- b = np.array([1.0,2.0])
- return np.dot(x,b) # The same as x_1*b_1 + x_2*b_2
-
-f9_alternative_grad = grad(f9_alternative)
-
-x = np.array([3.0,0.0])
-
-print("The gradient of f9 is:",f9_alternative_grad(x))
-
-# The analytical gradient of the dot product of vectors x and b with two elements (x_1,x_2) and (b_1, b_2) respectively
-# w.r.t x is (b_1, b_2).
+print("The derivative of f8 is:",f8_grad(x))
Here, Autograd tells us that an 'ArrayBox' does not support item assignment. The item assignment is done when the program tries to assign x[2] to the value 3. However, Autograd has implemented the computation of the derivative such that this assignment is not possible.
@@ -417,7 +382,7 @@ x = np.a
-
The documentation recommends to avoid inplace operations such as
+a += b
-a -= b
-a*= b
-a /=b
+ import autograd.numpy as np
+from autograd import grad
+def f9(a): # Assume a is an array with 2 elements
+ b = np.array([1.0,2.0])
+ return a.dot(b)
+
+f9_grad = grad(f9)
+
+x = np.array([1.0,0.0])
+
+print("The derivative of f9 is:",f9_grad(x))
+
+Here we are told that the 'dot' function does not belong to Autograd's +version of a Numpy array. To overcome this, an alternative syntax +which also computed the dot product can be used: +
+ + + +import autograd.numpy as np
+from autograd import grad
+def f9_alternative(x): # Assume a is an array with 2 elements
+ b = np.array([1.0,2.0])
+ return np.dot(x,b) # The same as x_1*b_1 + x_2*b_2
+
+f9_alternative_grad = grad(f9_alternative)
+
+x = np.array([3.0,0.0])
+
+print("The gradient of f9 is:",f9_alternative_grad(x))
+
+# The analytical gradient of the dot product of vectors x and b with two elements (x_1,x_2) and (b_1, b_2) respectively
+# w.r.t x is (b_1, b_2).
-
We conclude the part on optmization by showing how we can make codes -for linear regression and logistic regression using autograd. The -first example shows results with ordinary leats squares. -
- +The documentation recommends to avoid inplace operations such as
# Using Autograd to calculate gradients for OLS
-from random import random, seed
-import numpy as np
-import autograd.numpy as np
-import matplotlib.pyplot as plt
-from autograd import grad
-
-def CostOLS(beta):
- return (1.0/n)*np.sum((y-X @ beta)**2)
-
-n = 100
-x = 2*np.random.rand(n,1)
-y = 4+3*x+np.random.randn(n,1)
-
-X = np.c_[np.ones((n,1)), x]
-XT_X = X.T @ X
-theta_linreg = np.linalg.pinv(XT_X) @ (X.T @ y)
-print("Own inversion")
-print(theta_linreg)
-# Hessian matrix
-H = (2.0/n)* XT_X
-EigValues, EigVectors = np.linalg.eig(H)
-print(f"Eigenvalues of Hessian Matrix:{EigValues}")
-
-theta = np.random.randn(2,1)
-eta = 1.0/np.max(EigValues)
-Niterations = 1000
-# define the gradient
-training_gradient = grad(CostOLS)
-
-for iter in range(Niterations):
- gradients = training_gradient(theta)
- theta -= eta*gradients
-print("theta from own gd")
-print(theta)
-
-xnew = np.array([[0],[2]])
-Xnew = np.c_[np.ones((2,1)), xnew]
-ypredict = Xnew.dot(theta)
-ypredict2 = Xnew.dot(theta_linreg)
-
-plt.plot(xnew, ypredict, "r-")
-plt.plot(xnew, ypredict2, "b-")
-plt.plot(x, y ,'ro')
-plt.axis([0,2.0,0, 15.0])
-plt.xlabel(r'$x$')
-plt.ylabel(r'$y$')
-plt.title(r'Random numbers ')
-plt.show()
+ a += b
+a -= b
+a*= b
+a /=b
-
In this code we include the stochastic gradient descent approach discussed above. Note here that we specify which argument we are taking the derivative with respect to when using autograd.
+We conclude the part on optmization by showing how we can make codes +for linear regression and logistic regression using autograd. The +first example shows results with ordinary leats squares. +
@@ -326,17 +332,15 @@ MathJax.Hub.Config({# Using Autograd to calculate gradients using SGD
-# OLS example
+ # Using Autograd to calculate gradients for OLS
from random import random, seed
import numpy as np
import autograd.numpy as np
import matplotlib.pyplot as plt
from autograd import grad
-# Note change from previous example
-def CostOLS(y,X,theta):
- return np.sum((y-X @ theta)**2)
+def CostOLS(beta):
+ return (1.0/n)*np.sum((y-X @ beta)**2)
n = 100
x = 2*np.random.rand(n,1)
@@ -355,12 +359,11 @@ EigValues, EigVectors = np= np.random.randn(2,1)
eta = 1.0/np.max(EigValues)
Niterations = 1000
-
-# Note that we request the derivative wrt third argument (theta, 2 here)
-training_gradient = grad(CostOLS,2)
+# define the gradient
+training_gradient = grad(CostOLS)
for iter in range(Niterations):
- gradients = (1.0/n)*training_gradient(y, X, theta)
+ gradients = training_gradient(theta)
theta -= eta*gradients
print("theta from own gd")
print(theta)
@@ -378,27 +381,6 @@ plt.xlabel(r
plt.ylabel(r'$y$')
plt.title(r'Random numbers ')
plt.show()
-
-n_epochs = 50
-M = 5 #size of each minibatch
-m = int(n/M) #number of minibatches
-t0, t1 = 5, 50
-def learning_schedule(t):
- return t0/(t+t1)
-
-theta = np.random.randn(2,1)
-
-for epoch in range(n_epochs):
-# Can you figure out a better way of setting up the contributions to each batch?
- for i in range(m):
- random_index = M*np.random.randint(m)
- xi = X[random_index:random_index+M]
- yi = y[random_index:random_index+M]
- gradients = (2.0/M)*training_gradient(yi, xi, theta)
- eta = learning_schedule(epoch*m+i)
- theta = theta - eta*gradients
-print("theta from own sdg")
-print(theta)
-
In this code we include the stochastic gradient descent approach discussed above. Note here that we specify which argument we are taking the derivative with respect to when using autograd.
@@ -325,39 +328,79 @@ MathJax.Hub.Config({import autograd.numpy as np
+ # Using Autograd to calculate gradients using SGD
+# OLS example
+from random import random, seed
+import numpy as np
+import autograd.numpy as np
+import matplotlib.pyplot as plt
from autograd import grad
-def sigmoid(x):
- return 0.5 * (np.tanh(x / 2.) + 1)
+# Note change from previous example
+def CostOLS(y,X,theta):
+ return np.sum((y-X @ theta)**2)
-def logistic_predictions(weights, inputs):
- # Outputs probability of a label being true according to logistic model.
- return sigmoid(np.dot(inputs, weights))
+n = 100
+x = 2*np.random.rand(n,1)
+y = 4+3*x+np.random.randn(n,1)
-def training_loss(weights):
- # Training loss is the negative log-likelihood of the training labels.
- preds = logistic_predictions(weights, inputs)
- label_probabilities = preds * targets + (1 - preds) * (1 - targets)
- return -np.sum(np.log(label_probabilities))
+X = np.c_[np.ones((n,1)), x]
+XT_X = X.T @ X
+theta_linreg = np.linalg.pinv(XT_X) @ (X.T @ y)
+print("Own inversion")
+print(theta_linreg)
+# Hessian matrix
+H = (2.0/n)* XT_X
+EigValues, EigVectors = np.linalg.eig(H)
+print(f"Eigenvalues of Hessian Matrix:{EigValues}")
-# Build a toy dataset.
-inputs = np.array([[0.52, 1.12, 0.77],
- [0.88, -1.08, 0.15],
- [0.52, 0.06, -1.30],
- [0.74, -2.49, 1.39]])
-targets = np.array([True, True, False, True])
+theta = np.random.randn(2,1)
+eta = 1.0/np.max(EigValues)
+Niterations = 1000
-# Define a function that returns gradients of training loss using Autograd.
-training_gradient_fun = grad(training_loss)
+# Note that we request the derivative wrt third argument (theta, 2 here)
+training_gradient = grad(CostOLS,2)
-# Optimize weights using gradient descent.
-weights = np.array([0.0, 0.0, 0.0])
-print("Initial loss:", training_loss(weights))
-for i in range(100):
- weights -= training_gradient_fun(weights) * 0.01
+for iter in range(Niterations):
+ gradients = (1.0/n)*training_gradient(y, X, theta)
+ theta -= eta*gradients
+print("theta from own gd")
+print(theta)
-print("Trained loss:", training_loss(weights))
+xnew = np.array([[0],[2]])
+Xnew = np.c_[np.ones((2,1)), xnew]
+ypredict = Xnew.dot(theta)
+ypredict2 = Xnew.dot(theta_linreg)
+
+plt.plot(xnew, ypredict, "r-")
+plt.plot(xnew, ypredict2, "b-")
+plt.plot(x, y ,'ro')
+plt.axis([0,2.0,0, 15.0])
+plt.xlabel(r'$x$')
+plt.ylabel(r'$y$')
+plt.title(r'Random numbers ')
+plt.show()
+
+n_epochs = 50
+M = 5 #size of each minibatch
+m = int(n/M) #number of minibatches
+t0, t1 = 5, 50
+def learning_schedule(t):
+ return t0/(t+t1)
+
+theta = np.random.randn(2,1)
+
+for epoch in range(n_epochs):
+# Can you figure out a better way of setting up the contributions to each batch?
+ for i in range(m):
+ random_index = M*np.random.randint(m)
+ xi = X[random_index:random_index+M]
+ yi = y[random_index:random_index+M]
+ gradients = (2.0/M)*training_gradient(yi, xi, theta)
+ eta = learning_schedule(epoch*m+i)
+ theta = theta - eta*gradients
+print("theta from own sdg")
+print(theta)
-
import autograd.numpy as np
+from autograd import grad
+
+def sigmoid(x):
+ return 0.5 * (np.tanh(x / 2.) + 1)
+
+def logistic_predictions(weights, inputs):
+ # Outputs probability of a label being true according to logistic model.
+ return sigmoid(np.dot(inputs, weights))
+
+def training_loss(weights):
+ # Training loss is the negative log-likelihood of the training labels.
+ preds = logistic_predictions(weights, inputs)
+ label_probabilities = preds * targets + (1 - preds) * (1 - targets)
+ return -np.sum(np.log(label_probabilities))
+
+# Build a toy dataset.
+inputs = np.array([[0.52, 1.12, 0.77],
+ [0.88, -1.08, 0.15],
+ [0.52, 0.06, -1.30],
+ [0.74, -2.49, 1.39]])
+targets = np.array([True, True, False, True])
+
+# Define a function that returns gradients of training loss using Autograd.
+training_gradient_fun = grad(training_loss)
+
+# Optimize weights using gradient descent.
+weights = np.array([0.0, 0.0, 0.0])
+print("Initial loss:", training_loss(weights))
+for i in range(100):
+ weights -= training_gradient_fun(weights) * 0.01
+
+print("Trained loss:", training_loss(weights))
+
+@@ -347,7 +401,7 @@ MathJax.Hub.Config({
-
Artificial neural networks are computational systems that can learn to -perform tasks by considering examples, generally without being -programmed with any task-specific rules. It is supposed to mimic a -biological system, wherein neurons interact by sending signals in the -form of mathematical functions between layers. All layers can contain -an arbitrary number of neurons, and each connection is represented by -a weight variable. -
+Neural Networks demystified + +Building Neural Networks from scratch@@ -352,7 +349,7 @@ a weight variable.
-
The field of artificial neural networks has a long history of -development, and is closely connected with the advancement of computer -science and computers in general. A model of artificial neurons was -first developed by McCulloch and Pitts in 1943 to study signal -processing in the brain and has later been refined by others. The -general idea is to mimic neural networks in the human brain, which is -composed of billions of neurons that communicate with each other by -sending electrical signals. Each neuron accumulates its incoming -signals, which must exceed an activation threshold to yield an -output. If the threshold is not overcome, the neuron remains inactive, -i.e. has zero output. -
- -This behaviour has inspired a simple mathematical model for an artificial neuron.
- -$$ -\begin{equation} - y = f\left(\sum_{i=1}^n w_ix_i\right) = f(u) -\tag{6} -\end{equation} -$$ - -Here, the output \( y \) of the neuron is the value of its activation function, which have as input -a weighted sum of signals \( x_i, \dots ,x_n \) received by \( n \) other neurons. -
- -Conceptually, it is helpful to divide neural networks into four -categories: -
-In natural science, DNNs and CNNs have already found numerous -applications. In statistical physics, they have been applied to detect -phase transitions in 2D Ising and Potts models, lattice gauge -theories, and different phases of polymers, or solving the -Navier-Stokes equation in weather forecasting. Deep learning has also -found interesting applications in quantum physics. Various quantum -phase transitions can be detected and studied using DNNs and CNNs, -topological phases, and even non-equilibrium many-body -localization. Representing quantum states as DNNs quantum state -tomography are among some of the impressive achievements to reveal the -potential of DNNs to facilitate the study of quantum systems. -
- -In quantum information theory, it has been shown that one can perform -gate decompositions with the help of neural. -
- -The applications are not limited to the natural sciences. There is a -plethora of applications in essentially all disciplines, from the -humanities to life science and medicine. +
Artificial neural networks are computational systems that can learn to +perform tasks by considering examples, generally without being +programmed with any task-specific rules. It is supposed to mimic a +biological system, wherein neurons interact by sending signals in the +form of mathematical functions between layers. All layers can contain +an arbitrary number of neurons, and each connection is represented by +a weight variable.
@@ -400,7 +354,7 @@ humanities to life science and medicine.
-
An artificial neural network (ANN), is a computational model that -consists of layers of connected neurons, or nodes or units. We will -refer to these interchangeably as units or nodes, and sometimes as -neurons. +
The field of artificial neural networks has a long history of +development, and is closely connected with the advancement of computer +science and computers in general. A model of artificial neurons was +first developed by McCulloch and Pitts in 1943 to study signal +processing in the brain and has later been refined by others. The +general idea is to mimic neural networks in the human brain, which is +composed of billions of neurons that communicate with each other by +sending electrical signals. Each neuron accumulates its incoming +signals, which must exceed an activation threshold to yield an +output. If the threshold is not overcome, the neuron remains inactive, +i.e. has zero output.
-It is supposed to mimic a biological nervous system by letting each -neuron interact with other neurons by sending signals in the form of -mathematical functions between layers. A wide variety of different -ANNs have been developed, but most of them consist of an input layer, -an output layer and eventual layers in-between, called hidden -layers. All layers can contain an arbitrary number of nodes, and each -connection between two nodes is associated with a weight variable. +
This behaviour has inspired a simple mathematical model for an artificial neuron.
+ +$$ +\begin{equation} + y = f\left(\sum_{i=1}^n w_ix_i\right) = f(u) +\tag{6} +\end{equation} +$$ + +Here, the output \( y \) of the neuron is the value of its activation function, which have as input +a weighted sum of signals \( x_i, \dots ,x_n \) received by \( n \) other neurons.
-Neural networks (also called neural nets) are neural-inspired -nonlinear models for supervised learning. As we will see, neural nets -can be viewed as natural, more powerful extensions of supervised -learning methods such as linear and logistic regression and soft-max -methods we discussed earlier. +
Conceptually, it is helpful to divide neural networks into four +categories: +
+In natural science, DNNs and CNNs have already found numerous +applications. In statistical physics, they have been applied to detect +phase transitions in 2D Ising and Potts models, lattice gauge +theories, and different phases of polymers, or solving the +Navier-Stokes equation in weather forecasting. Deep learning has also +found interesting applications in quantum physics. Various quantum +phase transitions can be detected and studied using DNNs and CNNs, +topological phases, and even non-equilibrium many-body +localization. Representing quantum states as DNNs quantum state +tomography are among some of the impressive achievements to reveal the +potential of DNNs to facilitate the study of quantum systems. +
+ +In quantum information theory, it has been shown that one can perform +gate decompositions with the help of neural. +
+ +The applications are not limited to the natural sciences. There is a +plethora of applications in essentially all disciplines, from the +humanities to life science and medicine.
@@ -365,7 +402,7 @@ methods we discussed earlier.
-
The feed-forward neural network (FFNN) was the first and simplest type -of ANNs that were devised. In this network, the information moves in -only one direction: forward through the layers. +
An artificial neural network (ANN), is a computational model that +consists of layers of connected neurons, or nodes or units. We will +refer to these interchangeably as units or nodes, and sometimes as +neurons.
-Nodes are represented by circles, while the arrows display the -connections between the nodes, including the direction of information -flow. Additionally, each arrow corresponds to a weight variable -(figure to come). We observe that each node in a layer is connected -to all nodes in the subsequent layer, making this a so-called -fully-connected FFNN. +
It is supposed to mimic a biological nervous system by letting each +neuron interact with other neurons by sending signals in the form of +mathematical functions between layers. A wide variety of different +ANNs have been developed, but most of them consist of an input layer, +an output layer and eventual layers in-between, called hidden +layers. All layers can contain an arbitrary number of nodes, and each +connection between two nodes is associated with a weight variable. +
+ +Neural networks (also called neural nets) are neural-inspired +nonlinear models for supervised learning. As we will see, neural nets +can be viewed as natural, more powerful extensions of supervised +learning methods such as linear and logistic regression and soft-max +methods we discussed earlier.
@@ -356,7 +367,7 @@ to all nodes in the subsequent layer, making this a so-called
-
A different variant of FFNNs are convolutional neural networks -(CNNs), which have a connectivity pattern inspired by the animal -visual cortex. Individual neurons in the visual cortex only respond to -stimuli from small sub-regions of the visual field, called a receptive -field. This makes the neurons well-suited to exploit the strong -spatially local correlation present in natural images. The response of -each neuron can be approximated mathematically as a convolution -operation. (figure to come) +
The feed-forward neural network (FFNN) was the first and simplest type +of ANNs that were devised. In this network, the information moves in +only one direction: forward through the layers.
-Convolutional neural networks emulate the behaviour of neurons in the -visual cortex by enforcing a local connectivity pattern between -nodes of adjacent layers: Each node in a convolutional layer is -connected only to a subset of the nodes in the previous layer, in -contrast to the fully-connected FFNN. Often, CNNs consist of several -convolutional layers that learn local features of the input, with a -fully-connected layer at the end, which gathers all the local data and -produces the outputs. They have wide applications in image and video -recognition. +
Nodes are represented by circles, while the arrows display the +connections between the nodes, including the direction of information +flow. Additionally, each arrow corresponds to a weight variable +(figure to come). We observe that each node in a layer is connected +to all nodes in the subsequent layer, making this a so-called +fully-connected FFNN.
@@ -364,7 +358,7 @@ recognition.
-
So far we have only mentioned ANNs where information flows in one -direction: forward. Recurrent neural networks on the other hand, -have connections between nodes that form directed cycles. This -creates a form of internal memory which are able to capture -information on what has been calculated before; the output is -dependent on the previous computations. Recurrent NNs make use of -sequential information by performing the same task for every element -in a sequence, where each element depends on previous elements. An -example of such information is sentences, making recurrent NNs -especially well-suited for handwriting and speech recognition. +
A different variant of FFNNs are convolutional neural networks +(CNNs), which have a connectivity pattern inspired by the animal +visual cortex. Individual neurons in the visual cortex only respond to +stimuli from small sub-regions of the visual field, called a receptive +field. This makes the neurons well-suited to exploit the strong +spatially local correlation present in natural images. The response of +each neuron can be approximated mathematically as a convolution +operation. (figure to come) +
+ +Convolutional neural networks emulate the behaviour of neurons in the +visual cortex by enforcing a local connectivity pattern between +nodes of adjacent layers: Each node in a convolutional layer is +connected only to a subset of the nodes in the previous layer, in +contrast to the fully-connected FFNN. Often, CNNs consist of several +convolutional layers that learn local features of the input, with a +fully-connected layer at the end, which gathers all the local data and +produces the outputs. They have wide applications in image and video +recognition.
@@ -355,7 +366,7 @@ especially well-suited for handwriting and speech recognition.
-
There are many other kinds of ANNs that have been developed. One type -that is specifically designed for interpolation in multidimensional -space is the radial basis function (RBF) network. RBFs are typically -made up of three layers: an input layer, a hidden layer with -non-linear radial symmetric activation functions and a linear output -layer (''linear'' here means that each node in the output layer has a -linear activation function). The layers are normally fully-connected -and there are no cycles, thus RBFs can be viewed as a type of -fully-connected FFNN. They are however usually treated as a separate -type of NN due the unusual activation functions. +
So far we have only mentioned ANNs where information flows in one +direction: forward. Recurrent neural networks on the other hand, +have connections between nodes that form directed cycles. This +creates a form of internal memory which are able to capture +information on what has been calculated before; the output is +dependent on the previous computations. Recurrent NNs make use of +sequential information by performing the same task for every element +in a sequence, where each element depends on previous elements. An +example of such information is sentences, making recurrent NNs +especially well-suited for handwriting and speech recognition.
@@ -355,7 +357,7 @@ type of NN due the unusual activation functions.
-
One uses often so-called fully-connected feed-forward neural networks -with three or more layers (an input layer, one or more hidden layers -and an output layer) consisting of neurons that have non-linear -activation functions. +
There are many other kinds of ANNs that have been developed. One type +that is specifically designed for interpolation in multidimensional +space is the radial basis function (RBF) network. RBFs are typically +made up of three layers: an input layer, a hidden layer with +non-linear radial symmetric activation functions and a linear output +layer (''linear'' here means that each node in the output layer has a +linear activation function). The layers are normally fully-connected +and there are no cycles, thus RBFs can be viewed as a type of +fully-connected FFNN. They are however usually treated as a separate +type of NN due the unusual activation functions.
-Such networks are often called multilayer perceptrons (MLPs).
-diff --git a/doc/pub/week40/html/._week40-bs043.html b/doc/pub/week40/html/._week40-bs043.html index 5f14f3c55..762d5b233 100644 --- a/doc/pub/week40/html/._week40-bs043.html +++ b/doc/pub/week40/html/._week40-bs043.html @@ -63,6 +63,7 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d 2, None, 'program-for-stochastic-gradient'), + ('Replace or not', 2, None, 'replace-or-not'), ('Momentum based GD', 2, None, 'momentum-based-gd'), ('More on momentum based approaches', 2, @@ -249,62 +250,63 @@ MathJax.Hub.Config({
-
According to the Universal approximation theorem, a feed-forward -neural network with just a single hidden layer containing a finite -number of neurons can approximate a continuous multidimensional -function to arbitrary accuracy, assuming the activation function for -the hidden layer is a non-constant, bounded and -monotonically-increasing continuous function. +
One uses often so-called fully-connected feed-forward neural networks +with three or more layers (an input layer, one or more hidden layers +and an output layer) consisting of neurons that have non-linear +activation functions.
-Note that the requirements on the activation function only applies to -the hidden layer, the output nodes are always assumed to be linear, so -as to not restrict the range of output values. -
+Such networks are often called multilayer perceptrons (MLPs).
@@ -356,7 +353,7 @@ as to not restrict the range of output values.
-
Figure 1: In a) we show a single perceptron model while in b) we dispay a network with two hidden layers, an input layer and an output layer.
-
According to the Universal approximation theorem, a feed-forward +neural network with just a single hidden layer containing a finite +number of neurons can approximate a continuous multidimensional +function to arbitrary accuracy, assuming the activation function for +the hidden layer is a non-constant, bounded and +monotonically-increasing continuous function. +
+ +Note that the requirements on the activation function only applies to +the hidden layer, the output nodes are always assumed to be linear, so +as to not restrict the range of output values. +
@@ -351,7 +358,7 @@ MathJax.Hub.Config({
-
Let us first try to fit various gates using standard linear -regression. The gates we are thinking of are the classical XOR, OR and -AND gates, well-known elements in computer science. The tables here -show how we can set up the inputs \( x_1 \) and \( x_2 \) in order to yield a -specific target \( y_i \). -
- - - -"""
-Simple code that tests XOR, OR and AND gates with linear regression
-"""
-
-import numpy as np
-# Design matrix
-X = np.array([ [1, 0, 0], [1, 0, 1], [1, 1, 0],[1, 1, 1]],dtype=np.float64)
-print(f"The X.TX matrix:{X.T @ X}")
-Xinv = np.linalg.pinv(X.T @ X)
-print(f"The invers of X.TX matrix:{Xinv}")
-
-# The XOR gate
-yXOR = np.array( [ 0, 1 ,1, 0])
-ThetaXOR = Xinv @ X.T @ yXOR
-print(f"The values of theta for the XOR gate:{ThetaXOR}")
-print(f"The linear regression prediction for the XOR gate:{X @ ThetaXOR}")
-
-
-# The OR gate
-yOR = np.array( [ 0, 1 ,1, 1])
-ThetaOR = Xinv @ X.T @ yOR
-print(f"The values of theta for the OR gate:{ThetaOR}")
-print(f"The linear regression prediction for the OR gate:{X @ ThetaOR}")
-
-
-# The OR gate
-yAND = np.array( [ 0, 0 ,0, 1])
-ThetaAND = Xinv @ X.T @ yAND
-print(f"The values of theta for the AND gate:{ThetaAND}")
-print(f"The linear regression prediction for the AND gate:{X @ ThetaAND}")
-
-What is happening here?
+Figure 1: In a) we show a single perceptron model while in b) we dispay a network with two hidden layers, an input layer and an output layer.
+
@@ -404,7 +353,7 @@ ThetaAND = Xinv 54
Let us first try to fit various gates using standard linear
+regression. The gates we are thinking of are the classical XOR, OR and
+AND gates, well-known elements in computer science. The tables here
+show how we can set up the inputs \( x_1 \) and \( x_2 \) in order to yield a
+specific target \( y_i \).
+Does Logistic Regression do a better Job?
+Examples of XOR, OR and AND gates
+
+"""
-Simple code that tests XOR and OR gates with linear regression
-and logistic regression
+Simple code that tests XOR, OR and AND gates with linear regression
"""
-import matplotlib.pyplot as plt
-from sklearn.linear_model import LogisticRegression
import numpy as np
-
# Design matrix
X = np.array([ [1, 0, 0], [1, 0, 1], [1, 1, 0],[1, 1, 1]],dtype=np.float64)
print(f"The X.TX matrix:{X.T @ X}")
@@ -359,21 +364,6 @@ yAND = np.= Xinv @ X.T @ yAND
print(f"The values of theta for the AND gate:{ThetaAND}")
print(f"The linear regression prediction for the AND gate:{X @ ThetaAND}")
-
-# Now we change to logistic regression
-
-
-# Logistic Regression
-logreg = LogisticRegression()
-logreg.fit(X, yOR)
-print("Test set accuracy with Logistic Regression for OR gate: {:.2f}".format(logreg.score(X,yOR)))
-
-logreg.fit(X, yXOR)
-print("Test set accuracy with Logistic Regression for XOR gate: {:.2f}".format(logreg.score(X,yXOR)))
-
-
-logreg.fit(X, yAND)
-print("Test set accuracy with Logistic Regression for AND gate: {:.2f}".format(logreg.score(X,yAND)))
Not exactly impressive, but somewhat better.
+What is happening here?
@@ -416,7 +406,7 @@ logreg.fit(X, yAND)
-
# and now neural networks with Scikit-Learn and the XOR
+ """
+Simple code that tests XOR and OR gates with linear regression
+and logistic regression
+"""
-from sklearn.neural_network import MLPClassifier
-from sklearn.datasets import make_classification
-X, yXOR = make_classification(n_samples=100, random_state=1)
-FFNN = MLPClassifier(random_state=1, max_iter=300).fit(X, yXOR)
-FFNN.predict_proba(X)
-print(f"Test set accuracy with Feed Forward Neural Network for XOR gate:{FFNN.score(X, yXOR)}")
+import matplotlib.pyplot as plt
+from sklearn.linear_model import LogisticRegression
+import numpy as np
+
+# Design matrix
+X = np.array([ [1, 0, 0], [1, 0, 1], [1, 1, 0],[1, 1, 1]],dtype=np.float64)
+print(f"The X.TX matrix:{X.T @ X}")
+Xinv = np.linalg.pinv(X.T @ X)
+print(f"The invers of X.TX matrix:{Xinv}")
+
+# The XOR gate
+yXOR = np.array( [ 0, 1 ,1, 0])
+ThetaXOR = Xinv @ X.T @ yXOR
+print(f"The values of theta for the XOR gate:{ThetaXOR}")
+print(f"The linear regression prediction for the XOR gate:{X @ ThetaXOR}")
+
+
+# The OR gate
+yOR = np.array( [ 0, 1 ,1, 1])
+ThetaOR = Xinv @ X.T @ yOR
+print(f"The values of theta for the OR gate:{ThetaOR}")
+print(f"The linear regression prediction for the OR gate:{X @ ThetaOR}")
+
+
+# The OR gate
+yAND = np.array( [ 0, 0 ,0, 1])
+ThetaAND = Xinv @ X.T @ yAND
+print(f"The values of theta for the AND gate:{ThetaAND}")
+print(f"The linear regression prediction for the AND gate:{X @ ThetaAND}")
+
+# Now we change to logistic regression
+
+
+# Logistic Regression
+logreg = LogisticRegression()
+logreg.fit(X, yOR)
+print("Test set accuracy with Logistic Regression for OR gate: {:.2f}".format(logreg.score(X,yOR)))
+
+logreg.fit(X, yXOR)
+print("Test set accuracy with Logistic Regression for XOR gate: {:.2f}".format(logreg.score(X,yXOR)))
+
+
+logreg.fit(X, yAND)
+print("Test set accuracy with Logistic Regression for AND gate: {:.2f}".format(logreg.score(X,yAND)))
Not exactly impressive, but somewhat better.
@@ -374,7 +418,7 @@ FFNN.predict_proba(X)
-
The output \( y \) is produced via the activation function \( f \)
-$$ - y = f\left(\sum_{i=1}^n w_ix_i + b_i\right) = f(z), -$$ -This function receives \( x_i \) as inputs. -Here the activation \( z=(\sum_{i=1}^n w_ix_i+b_i) \). -In an FFNN of such neurons, the inputs \( x_i \) are the outputs of -the neurons in the preceding layer. Furthermore, an MLP is -fully-connected, which means that each neuron receives a weighted sum -of the outputs of all neurons in the previous layer. -
+ +# and now neural networks with Scikit-Learn and the XOR
+
+from sklearn.neural_network import MLPClassifier
+from sklearn.datasets import make_classification
+X, yXOR = make_classification(n_samples=100, random_state=1)
+FFNN = MLPClassifier(random_state=1, max_iter=300).fit(X, yXOR)
+FFNN.predict_proba(X)
+print(f"Test set accuracy with Feed Forward Neural Network for XOR gate:{FFNN.score(X, yXOR)}")
+
+@@ -356,7 +376,7 @@ of the outputs of all neurons in the previous layer.
First, for each node \( i \) in the first hidden layer, we calculate a weighted sum \( z_i^1 \) of the input coordinates \( x_j \),
- +The output \( y \) is produced via the activation function \( f \)
$$ -\begin{equation} z_i^1 = \sum_{j=1}^{M} w_{ij}^1 x_j + b_i^1 -\tag{7} -\end{equation} + y = f\left(\sum_{i=1}^n w_ix_i + b_i\right) = f(z), $$ -Here \( b_i \) is the so-called bias which is normally needed in -case of zero activation weights or inputs. How to fix the biases and -the weights will be discussed below. The value of \( z_i^1 \) is the -argument to the activation function \( f_i \) of each node \( i \), The -variable \( M \) stands for all possible inputs to a given node \( i \) in the -first layer. We define the output \( y_i^1 \) of all neurons in layer 1 as -
- -$$ -\begin{equation} - y_i^1 = f(z_i^1) = f\left(\sum_{j=1}^M w_{ij}^1 x_j + b_i^1\right) -\tag{8} -\end{equation} -$$ - -where we assume that all nodes in the same layer have identical -activation functions, hence the notation \( f \). In general, we could assume in the more general case that different layers have different activation functions. -In this case we would identify these functions with a superscript \( l \) for the \( l \)-th layer, -
- -$$ -\begin{equation} - y_i^l = f^l(u_i^l) = f^l\left(\sum_{j=1}^{N_{l-1}} w_{ij}^l y_j^{l-1} + b_i^l\right) -\tag{9} -\end{equation} -$$ - -where \( N_l \) is the number of nodes in layer \( l \). When the output of -all the nodes in the first hidden layer are computed, the values of -the subsequent layer can be calculated and so forth until the output -is obtained. +
This function receives \( x_i \) as inputs. +Here the activation \( z=(\sum_{i=1}^n w_ix_i+b_i) \). +In an FFNN of such neurons, the inputs \( x_i \) are the outputs of +the neurons in the preceding layer. Furthermore, an MLP is +fully-connected, which means that each neuron receives a weighted sum +of the outputs of all neurons in the previous layer.
@@ -384,7 +358,7 @@ is obtained.
The output of neuron \( i \) in layer 2 is thus,
+First, for each node \( i \) in the first hidden layer, we calculate a weighted sum \( z_i^1 \) of the input coordinates \( x_j \),
$$ -\begin{align} - y_i^2 &= f^2\left(\sum_{j=1}^N w_{ij}^2 y_j^1 + b_i^2\right) -\tag{10}\\ - &= f^2\left[\sum_{j=1}^N w_{ij}^2f^1\left(\sum_{k=1}^M w_{jk}^1 x_k + b_j^1\right) + b_i^2\right] -\tag{11} -\end{align} +\begin{equation} z_i^1 = \sum_{j=1}^{M} w_{ij}^1 x_j + b_i^1 +\tag{7} +\end{equation} $$ -where we have substituted \( y_k^1 \) with the inputs \( x_k \). Finally, the ANN output reads
+Here \( b_i \) is the so-called bias which is normally needed in +case of zero activation weights or inputs. How to fix the biases and +the weights will be discussed below. The value of \( z_i^1 \) is the +argument to the activation function \( f_i \) of each node \( i \), The +variable \( M \) stands for all possible inputs to a given node \( i \) in the +first layer. We define the output \( y_i^1 \) of all neurons in layer 1 as +
$$ -\begin{align} - y_i^3 &= f^3\left(\sum_{j=1}^N w_{ij}^3 y_j^2 + b_i^3\right) -\tag{12}\\ - &= f_3\left[\sum_{j} w_{ij}^3 f^2\left(\sum_{k} w_{jk}^2 f^1\left(\sum_{m} w_{km}^1 x_m + b_k^1\right) + b_j^2\right) - + b_1^3\right] -\tag{13} -\end{align} +\begin{equation} + y_i^1 = f(z_i^1) = f\left(\sum_{j=1}^M w_{ij}^1 x_j + b_i^1\right) +\tag{8} +\end{equation} $$ +where we assume that all nodes in the same layer have identical +activation functions, hence the notation \( f \). In general, we could assume in the more general case that different layers have different activation functions. +In this case we would identify these functions with a superscript \( l \) for the \( l \)-th layer, +
+ +$$ +\begin{equation} + y_i^l = f^l(u_i^l) = f^l\left(\sum_{j=1}^{N_{l-1}} w_{ij}^l y_j^{l-1} + b_i^l\right) +\tag{9} +\end{equation} +$$ + +where \( N_l \) is the number of nodes in layer \( l \). When the output of +all the nodes in the first hidden layer are computed, the values of +the subsequent layer can be calculated and so forth until the output +is obtained. +
@@ -367,7 +386,7 @@ $$
We can generalize this expression to an MLP with \( l \) hidden -layers. The complete functional form is, -
+The output of neuron \( i \) in layer 2 is thus,
$$ \begin{align} -&y^{l+1}_i = f^{l+1}\left[\!\sum_{j=1}^{N_l} w_{ij}^3 f^l\left(\sum_{k=1}^{N_{l-1}}w_{jk}^{l-1}\left(\dots f^1\left(\sum_{n=1}^{N_0} w_{mn}^1 x_n+ b_m^1\right)\dots\right)+b_k^2\right)+b_1^3\right] && -\tag{14} + y_i^2 &= f^2\left(\sum_{j=1}^N w_{ij}^2 y_j^1 + b_i^2\right) +\tag{10}\\ + &= f^2\left[\sum_{j=1}^N w_{ij}^2f^1\left(\sum_{k=1}^M w_{jk}^1 x_k + b_j^1\right) + b_i^2\right] +\tag{11} +\end{align} +$$ + +where we have substituted \( y_k^1 \) with the inputs \( x_k \). Finally, the ANN output reads
+ +$$ +\begin{align} + y_i^3 &= f^3\left(\sum_{j=1}^N w_{ij}^3 y_j^2 + b_i^3\right) +\tag{12}\\ + &= f_3\left[\sum_{j} w_{ij}^3 f^2\left(\sum_{k} w_{jk}^2 f^1\left(\sum_{m} w_{km}^1 x_m + b_k^1\right) + b_j^2\right) + + b_1^3\right] +\tag{13} \end{align} $$ -which illustrates a basic property of MLPs: The only independent -variables are the input values \( x_n \). -
@@ -358,7 +369,7 @@ variables are the input values \( x_n \).
This confirms that an MLP, despite its quite convoluted mathematical -form, is nothing more than an analytic function, specifically a -mapping of real-valued vectors \( \hat{x} \in \mathbb{R}^n \rightarrow -\hat{y} \in \mathbb{R}^m \). -
- -Furthermore, the flexibility and universality of an MLP can be -illustrated by realizing that the expression is essentially a nested -sum of scaled activation functions of the form +
We can generalize this expression to an MLP with \( l \) hidden +layers. The complete functional form is,
$$ -\begin{equation} - f(x) = c_1 f(c_2 x + c_3) + c_4 -\tag{15} -\end{equation} +\begin{align} +&y^{l+1}_i = f^{l+1}\left[\!\sum_{j=1}^{N_l} w_{ij}^3 f^l\left(\sum_{k=1}^{N_{l-1}}w_{jk}^{l-1}\left(\dots f^1\left(\sum_{n=1}^{N_0} w_{mn}^1 x_n+ b_m^1\right)\dots\right)+b_k^2\right)+b_1^3\right] && +\tag{14} +\end{align} $$ -where the parameters \( c_i \) are weights and biases. By adjusting these -parameters, the activation functions can be shifted up and down or -left and right, change slope or be rescaled which is the key to the -flexibility of a neural network. +
which illustrates a basic property of MLPs: The only independent +variables are the input values \( x_n \).
@@ -367,7 +360,7 @@ flexibility of a neural network.
-
We can introduce a more convenient notation for the activations in an A NN.
- -Additionally, we can represent the biases and activations -as layer-wise column vectors \( \hat{b}_l \) and \( \hat{y}_l \), so that the \( i \)-th element of each vector -is the bias \( b_i^l \) and activation \( y_i^l \) of node \( i \) in layer \( l \) respectively. +
This confirms that an MLP, despite its quite convoluted mathematical +form, is nothing more than an analytic function, specifically a +mapping of real-valued vectors \( \hat{x} \in \mathbb{R}^n \rightarrow +\hat{y} \in \mathbb{R}^m \).
-We have that \( \mathrm{W}_l \) is an \( N_{l-1} \times N_l \) matrix, while \( \hat{b}_l \) and \( \hat{y}_l \) are \( N_l \times 1 \) column vectors. -With this notation, the sum becomes a matrix-vector multiplication, and we can write -the equation for the activations of hidden layer 2 (assuming three nodes for simplicity) as +
Furthermore, the flexibility and universality of an MLP can be +illustrated by realizing that the expression is essentially a nested +sum of scaled activation functions of the form
+ $$ \begin{equation} - \hat{y}_2 = f_2(\mathrm{W}_2 \hat{y}_{1} + \hat{b}_{2}) = - f_2\left(\left[\begin{array}{ccc} - w^2_{11} &w^2_{12} &w^2_{13} \\ - w^2_{21} &w^2_{22} &w^2_{23} \\ - w^2_{31} &w^2_{32} &w^2_{33} \\ - \end{array} \right] \cdot - \left[\begin{array}{c} - y^1_1 \\ - y^1_2 \\ - y^1_3 \\ - \end{array}\right] + - \left[\begin{array}{c} - b^2_1 \\ - b^2_2 \\ - b^2_3 \\ - \end{array}\right]\right). -\tag{16} + f(x) = c_1 f(c_2 x + c_3) + c_4 +\tag{15} \end{equation} $$ +where the parameters \( c_i \) are weights and biases. By adjusting these +parameters, the activation functions can be shifted up and down or +left and right, change slope or be rescaled which is the key to the +flexibility of a neural network. +
@@ -377,7 +369,7 @@ $$
-
The activation of node \( i \) in layer 2 is
+We can introduce a more convenient notation for the activations in an A NN.
+Additionally, we can represent the biases and activations +as layer-wise column vectors \( \hat{b}_l \) and \( \hat{y}_l \), so that the \( i \)-th element of each vector +is the bias \( b_i^l \) and activation \( y_i^l \) of node \( i \) in layer \( l \) respectively. +
+ +We have that \( \mathrm{W}_l \) is an \( N_{l-1} \times N_l \) matrix, while \( \hat{b}_l \) and \( \hat{y}_l \) are \( N_l \times 1 \) column vectors. +With this notation, the sum becomes a matrix-vector multiplication, and we can write +the equation for the activations of hidden layer 2 (assuming three nodes for simplicity) as +
$$ \begin{equation} - y^2_i = f_2\Bigr(w^2_{i1}y^1_1 + w^2_{i2}y^1_2 + w^2_{i3}y^1_3 + b^2_i\Bigr) = - f_2\left(\sum_{j=1}^3 w^2_{ij} y_j^1 + b^2_i\right). -\tag{17} + \hat{y}_2 = f_2(\mathrm{W}_2 \hat{y}_{1} + \hat{b}_{2}) = + f_2\left(\left[\begin{array}{ccc} + w^2_{11} &w^2_{12} &w^2_{13} \\ + w^2_{21} &w^2_{22} &w^2_{23} \\ + w^2_{31} &w^2_{32} &w^2_{33} \\ + \end{array} \right] \cdot + \left[\begin{array}{c} + y^1_1 \\ + y^1_2 \\ + y^1_3 \\ + \end{array}\right] + + \left[\begin{array}{c} + b^2_1 \\ + b^2_2 \\ + b^2_3 \\ + \end{array}\right]\right). +\tag{16} \end{equation} $$ -This is not just a convenient and compact notation, but also a useful -and intuitive way to think about MLPs: The output is calculated by a -series of matrix-vector multiplications and vector additions that are -used as input to the activation functions. For each operation -\( \mathrm{W}_l \hat{y}_{l-1} \) we move forward one layer. -
@@ -360,7 +379,7 @@ used as input to the activation functions. For each operation
-
A property that characterizes a neural network, other than its -connectivity, is the choice of activation function(s). As described -in, the following restrictions are imposed on an activation function -for a FFNN to fulfill the universal approximation theorem +
The activation of node \( i \) in layer 2 is
+ +$$ +\begin{equation} + y^2_i = f_2\Bigr(w^2_{i1}y^1_1 + w^2_{i2}y^1_2 + w^2_{i3}y^1_3 + b^2_i\Bigr) = + f_2\left(\sum_{j=1}^3 w^2_{ij} y_j^1 + b^2_i\right). +\tag{17} +\end{equation} +$$ + +This is not just a convenient and compact notation, but also a useful +and intuitive way to think about MLPs: The output is calculated by a +series of matrix-vector multiplications and vector additions that are +used as input to the activation functions. For each operation +\( \mathrm{W}_l \hat{y}_{l-1} \) we move forward one layer.
-diff --git a/doc/pub/week40/html/._week40-bs056.html b/doc/pub/week40/html/._week40-bs056.html index 7331e07d2..ef656d85a 100644 --- a/doc/pub/week40/html/._week40-bs056.html +++ b/doc/pub/week40/html/._week40-bs056.html @@ -63,6 +63,7 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d 2, None, 'program-for-stochastic-gradient'), + ('Replace or not', 2, None, 'replace-or-not'), ('Momentum based GD', 2, None, 'momentum-based-gd'), ('More on momentum based approaches', 2, @@ -249,62 +250,63 @@ MathJax.Hub.Config({
-
The second requirement excludes all linear functions. Furthermore, in -a MLP with only linear activation functions, each layer simply -performs a linear transformation of its inputs. +
A property that characterizes a neural network, other than its +connectivity, is the choice of activation function(s). As described +in, the following restrictions are imposed on an activation function +for a FFNN to fulfill the universal approximation theorem
-Regardless of the number of layers, the output of the NN will be -nothing but a linear function of the inputs. Thus we need to introduce -some kind of non-linearity to the NN to be able to fit non-linear -functions Typical examples are the logistic Sigmoid -
- -$$ - f(x) = \frac{1}{1 + e^{-x}}, -$$ - -and the hyperbolic tangent function
-$$ - f(x) = \tanh(x) -$$ - - +diff --git a/doc/pub/week40/html/._week40-bs057.html b/doc/pub/week40/html/._week40-bs057.html index e3dedf15b..fc9e64d0f 100644 --- a/doc/pub/week40/html/._week40-bs057.html +++ b/doc/pub/week40/html/._week40-bs057.html @@ -63,6 +63,7 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d 2, None, 'program-for-stochastic-gradient'), + ('Replace or not', 2, None, 'replace-or-not'), ('Momentum based GD', 2, None, 'momentum-based-gd'), ('More on momentum based approaches', 2, @@ -249,62 +250,63 @@ MathJax.Hub.Config({
-
The sigmoid function are more biologically plausible because the -output of inactive neurons are zero. Such activation function are -called one-sided. However, it has been shown that the hyperbolic -tangent performs better than the sigmoid for training MLPs. has -become the most popular for deep neural networks +
The second requirement excludes all linear functions. Furthermore, in +a MLP with only linear activation functions, each layer simply +performs a linear transformation of its inputs.
+Regardless of the number of layers, the output of the NN will be +nothing but a linear function of the inputs. Thus we need to introduce +some kind of non-linearity to the NN to be able to fit non-linear +functions Typical examples are the logistic Sigmoid +
- -"""The sigmoid function (or the logistic curve) is a
-function that takes any real number, z, and outputs a number (0,1).
-It is useful in neural networks for assigning weights on a relative scale.
-The value z is the weighted sum of parameters involved in the learning algorithm."""
+$$
+ f(x) = \frac{1}{1 + e^{-x}},
+$$
-import numpy
-import matplotlib.pyplot as plt
-import math as mt
-
-z = numpy.arange(-5, 5, .1)
-sigma_fn = numpy.vectorize(lambda z: 1/(1+numpy.exp(-z)))
-sigma = sigma_fn(z)
-
-fig = plt.figure()
-ax = fig.add_subplot(111)
-ax.plot(z, sigma)
-ax.set_ylim([-0.1, 1.1])
-ax.set_xlim([-5,5])
-ax.grid(True)
-ax.set_xlabel('z')
-ax.set_title('sigmoid function')
-
-plt.show()
-
-"""Step Function"""
-z = numpy.arange(-5, 5, .02)
-step_fn = numpy.vectorize(lambda z: 1.0 if z >= 0.0 else 0.0)
-step = step_fn(z)
-
-fig = plt.figure()
-ax = fig.add_subplot(111)
-ax.plot(z, step)
-ax.set_ylim([-0.5, 1.5])
-ax.set_xlim([-5,5])
-ax.grid(True)
-ax.set_xlabel('z')
-ax.set_title('step function')
-
-plt.show()
-
-"""Sine Function"""
-z = numpy.arange(-2*mt.pi, 2*mt.pi, 0.1)
-t = numpy.sin(z)
-
-fig = plt.figure()
-ax = fig.add_subplot(111)
-ax.plot(z, t)
-ax.set_ylim([-1.0, 1.0])
-ax.set_xlim([-2*mt.pi,2*mt.pi])
-ax.grid(True)
-ax.set_xlabel('z')
-ax.set_title('sine function')
-
-plt.show()
-
-"""Plots a graph of the squashing function used by a rectified linear
-unit"""
-z = numpy.arange(-2, 2, .1)
-zero = numpy.zeros(len(z))
-y = numpy.max([zero, z], axis=0)
-
-fig = plt.figure()
-ax = fig.add_subplot(111)
-ax.plot(z, y)
-ax.set_ylim([-2.0, 2.0])
-ax.set_xlim([-2.0, 2.0])
-ax.grid(True)
-ax.set_xlabel('z')
-ax.set_title('Rectified linear unit')
-
-plt.show()
-
-and the hyperbolic tangent function
+$$ + f(x) = \tanh(x) +$$@@ -444,7 +366,7 @@ plt.show()
-
The multilayer perceptron is a very popular, and easy to implement approach, to deep learning. It consists of
-As a convention it is normal to call a network with one layer of input units, one layer of hidden -units and one layer of output units as a two-layer network. A network with two layers of hidden units is called a three-layer network etc etc. +
The sigmoid function are more biologically plausible because the +output of inactive neurons are zero. Such activation function are +called one-sided. However, it has been shown that the hyperbolic +tangent performs better than the sigmoid for training MLPs. has +become the most popular for deep neural networks
-For an MLP network there is no direct connection between the output nodes/neurons/units and the input nodes/neurons/units. -Hereafter we will call the various entities of a layer for nodes. -There are also no connections within a single layer. -
-The number of input nodes does not need to equal the number of output -nodes. This applies also to the hidden layers. Each layer may have its -own number of nodes and activation functions. -
+ +"""The sigmoid function (or the logistic curve) is a
+function that takes any real number, z, and outputs a number (0,1).
+It is useful in neural networks for assigning weights on a relative scale.
+The value z is the weighted sum of parameters involved in the learning algorithm."""
+
+import numpy
+import matplotlib.pyplot as plt
+import math as mt
+
+z = numpy.arange(-5, 5, .1)
+sigma_fn = numpy.vectorize(lambda z: 1/(1+numpy.exp(-z)))
+sigma = sigma_fn(z)
+
+fig = plt.figure()
+ax = fig.add_subplot(111)
+ax.plot(z, sigma)
+ax.set_ylim([-0.1, 1.1])
+ax.set_xlim([-5,5])
+ax.grid(True)
+ax.set_xlabel('z')
+ax.set_title('sigmoid function')
+
+plt.show()
+
+"""Step Function"""
+z = numpy.arange(-5, 5, .02)
+step_fn = numpy.vectorize(lambda z: 1.0 if z >= 0.0 else 0.0)
+step = step_fn(z)
+
+fig = plt.figure()
+ax = fig.add_subplot(111)
+ax.plot(z, step)
+ax.set_ylim([-0.5, 1.5])
+ax.set_xlim([-5,5])
+ax.grid(True)
+ax.set_xlabel('z')
+ax.set_title('step function')
+
+plt.show()
+
+"""Sine Function"""
+z = numpy.arange(-2*mt.pi, 2*mt.pi, 0.1)
+t = numpy.sin(z)
+
+fig = plt.figure()
+ax = fig.add_subplot(111)
+ax.plot(z, t)
+ax.set_ylim([-1.0, 1.0])
+ax.set_xlim([-2*mt.pi,2*mt.pi])
+ax.grid(True)
+ax.set_xlabel('z')
+ax.set_title('sine function')
+
+plt.show()
+
+"""Plots a graph of the squashing function used by a rectified linear
+unit"""
+z = numpy.arange(-2, 2, .1)
+zero = numpy.zeros(len(z))
+y = numpy.max([zero, z], axis=0)
+
+fig = plt.figure()
+ax = fig.add_subplot(111)
+ax.plot(z, y)
+ax.set_ylim([-2.0, 2.0])
+ax.set_xlim([-2.0, 2.0])
+ax.grid(True)
+ax.set_xlabel('z')
+ax.set_title('Rectified linear unit')
+
+plt.show()
+
+The hidden layers have their name from the fact that they are not -linked to observables and as we will see below when we define the -so-called activation \( \hat{z} \), we can think of this as a basis -expansion of the original inputs \( \hat{x} \). The difference however -between neural networks and say linear regression is that now these -basis functions (which will correspond to the weights in the network) -are learned from data. This results in an important difference between -neural networks and deep learning approaches on one side and methods -like logistic regression or linear regression and their modifications on the other side. -
@@ -374,7 +446,7 @@ like logistic regression or linear regression and their modifications on the oth
-
A neural network with only one layer, what we called the simple -perceptron, is best suited if we have a standard binary model with -clear (linear) boundaries between the outcomes. As such it could -equally well be replaced by standard linear regression or logistic -regression. Networks with one or more hidden layers approximate -systems with more complex boundaries. +
The multilayer perceptron is a very popular, and easy to implement approach, to deep learning. It consists of
+As a convention it is normal to call a network with one layer of input units, one layer of hidden +units and one layer of output units as a two-layer network. A network with two layers of hidden units is called a three-layer network etc etc.
-As stated earlier, -an important theorem in studies of neural networks, restated without -proof here, is the universal approximation -theorem. +
For an MLP network there is no direct connection between the output nodes/neurons/units and the input nodes/neurons/units. +Hereafter we will call the various entities of a layer for nodes. +There are also no connections within a single layer.
-It states that a feed-forward network with a single hidden layer -containing a finite number of neurons can approximate continuous -functions on compact subsets of real functions. The theorem thus -states that simple neural networks can represent a wide variety of -interesting functions when given appropriate parameters. It is the -multilayer feedforward architecture itself which gives neural networks -the potential of being universal approximators. +
The number of input nodes does not need to equal the number of output +nodes. This applies also to the hidden layers. Each layer may have its +own number of nodes and activation functions. +
+ +The hidden layers have their name from the fact that they are not +linked to observables and as we will see below when we define the +so-called activation \( \hat{z} \), we can think of this as a basis +expansion of the original inputs \( \hat{x} \). The difference however +between neural networks and say linear regression is that now these +basis functions (which will correspond to the weights in the network) +are learned from data. This results in an important difference between +neural networks and deep learning approaches on one side and methods +like logistic regression or linear regression and their modifications on the other side.
@@ -365,6 +375,8 @@ the potential of being universal approximators.
-
As we have seen now in a feed forward network, we can express the final output of our network in terms of basic matrix-vector multiplications. -The unknowwn quantities are our weights \( w_{ij} \) and we need to find an algorithm for changing them so that our errors are as small as possible. -This leads us to the famous back propagation algorithm. +
A neural network with only one layer, what we called the simple +perceptron, is best suited if we have a standard binary model with +clear (linear) boundaries between the outcomes. As such it could +equally well be replaced by standard linear regression or logistic +regression. Networks with one or more hidden layers approximate +systems with more complex boundaries.
-The questions we want to ask are how do changes in the biases and the -weights in our network change the cost function and how can we use the -final output to modify the weights? +
As stated earlier, +an important theorem in studies of neural networks, restated without +proof here, is the universal approximation +theorem.
-To derive these equations let us start with a plain regression problem -and define our cost function as -
- -$$ -{\cal C}(\hat{W}) = \frac{1}{2}\sum_{i=1}^n\left(y_i - t_i\right)^2, -$$ - -where the $t_i$s are our \( n \) targets (the values we want to -reproduce), while the outputs of the network after having propagated -all inputs \( \hat{x} \) are given by \( y_i \). Below we will demonstrate -how the basic equations arising from the back propagation algorithm -can be modified in order to study classification problems with \( K \) -classes. +
It states that a feed-forward network with a single hidden layer +containing a finite number of neurons can approximate continuous +functions on compact subsets of real functions. The theorem thus +states that simple neural networks can represent a wide variety of +interesting functions when given appropriate parameters. It is the +multilayer feedforward architecture itself which gives neural networks +the potential of being universal approximators.
@@ -367,6 +366,7 @@ classes.
-
With our definition of the targets \( \hat{t} \), the outputs of the -network \( \hat{y} \) and the inputs \( \hat{x} \) we -define now the activation \( z_j^l \) of node/neuron/unit \( j \) of the -\( l \)-th layer as a function of the bias, the weights which add up from -the previous layer \( l-1 \) and the forward passes/outputs -\( \hat{a}^{l-1} \) from the previous layer as +
As we have seen now in a feed forward network, we can express the final output of our network in terms of basic matrix-vector multiplications. +The unknowwn quantities are our weights \( w_{ij} \) and we need to find an algorithm for changing them so that our errors are as small as possible. +This leads us to the famous back propagation algorithm. +
+ +The questions we want to ask are how do changes in the biases and the +weights in our network change the cost function and how can we use the +final output to modify the weights? +
+ +To derive these equations let us start with a plain regression problem +and define our cost function as
$$ -z_j^l = \sum_{i=1}^{M_{l-1}}w_{ij}^la_i^{l-1}+b_j^l, +{\cal C}(\hat{W}) = \frac{1}{2}\sum_{i=1}^n\left(y_i - t_i\right)^2, $$ -where \( b_k^l \) are the biases from layer \( l \). Here \( M_{l-1} \) -represents the total number of nodes/neurons/units of layer \( l-1 \). The -figure here illustrates this equation. We can rewrite this in a more -compact form as the matrix-vector products we discussed earlier, +
where the $t_i$s are our \( n \) targets (the values we want to +reproduce), while the outputs of the network after having propagated +all inputs \( \hat{x} \) are given by \( y_i \). Below we will demonstrate +how the basic equations arising from the back propagation algorithm +can be modified in order to study classification problems with \( K \) +classes.
-$$ -\hat{z}^l = \left(\hat{W}^l\right)^T\hat{a}^{l-1}+\hat{b}^l. -$$ - -With the activation values \( \hat{z}^l \) we can in turn define the -output of layer \( l \) as \( \hat{a}^l = f(\hat{z}^l) \) where \( f \) is our -activation function. In the examples here we will use the sigmoid -function discussed in our logistic regression lectures. We will also use the same activation function \( f \) for all layers -and their nodes. It means we have -
- -$$ -a_j^l = f(z_j^l) = \frac{1}{1+\exp{-(z_j^l)}}. -$$ - -diff --git a/doc/pub/week40/html/._week40-bs062.html b/doc/pub/week40/html/._week40-bs062.html index bbefd0b39..0ca7615b2 100644 --- a/doc/pub/week40/html/._week40-bs062.html +++ b/doc/pub/week40/html/._week40-bs062.html @@ -63,6 +63,7 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d 2, None, 'program-for-stochastic-gradient'), + ('Replace or not', 2, None, 'replace-or-not'), ('Momentum based GD', 2, None, 'momentum-based-gd'), ('More on momentum based approaches', 2, @@ -249,62 +250,63 @@ MathJax.Hub.Config({
-
With our definition of the targets \( \hat{t} \), the outputs of the +network \( \hat{y} \) and the inputs \( \hat{x} \) we +define now the activation \( z_j^l \) of node/neuron/unit \( j \) of the +\( l \)-th layer as a function of the bias, the weights which add up from +the previous layer \( l-1 \) and the forward passes/outputs +\( \hat{a}^{l-1} \) from the previous layer as +
-From the definition of the activation \( z_j^l \) we have
$$ -\frac{\partial z_j^l}{\partial w_{ij}^l} = a_i^{l-1}, +z_j^l = \sum_{i=1}^{M_{l-1}}w_{ij}^la_i^{l-1}+b_j^l, $$ -and
+where \( b_k^l \) are the biases from layer \( l \). Here \( M_{l-1} \) +represents the total number of nodes/neurons/units of layer \( l-1 \). The +figure here illustrates this equation. We can rewrite this in a more +compact form as the matrix-vector products we discussed earlier, +
+ $$ -\frac{\partial z_j^l}{\partial a_i^{l-1}} = w_{ji}^l. +\hat{z}^l = \left(\hat{W}^l\right)^T\hat{a}^{l-1}+\hat{b}^l. $$ -With our definition of the activation function we have that (note that this function depends only on \( z_j^l \))
+With the activation values \( \hat{z}^l \) we can in turn define the +output of layer \( l \) as \( \hat{a}^l = f(\hat{z}^l) \) where \( f \) is our +activation function. In the examples here we will use the sigmoid +function discussed in our logistic regression lectures. We will also use the same activation function \( f \) for all layers +and their nodes. It means we have +
+ $$ -\frac{\partial a_j^l}{\partial z_j^{l}} = a_j^l(1-a_j^l)=f(z_j^l)(1-f(z_j^l)). +a_j^l = f(z_j^l) = \frac{1}{1+\exp{-(z_j^l)}}. $$ @@ -355,6 +375,7 @@ $$
-
With these definitions we can now compute the derivative of the cost function in terms of the weights.
- -Let us specialize to the output layer \( l=L \). Our cost function is
+From the definition of the activation \( z_j^l \) we have
$$ -{\cal C}(\hat{W^L}) = \frac{1}{2}\sum_{i=1}^n\left(y_i - t_i\right)^2=\frac{1}{2}\sum_{i=1}^n\left(a_i^L - t_i\right)^2, +\frac{\partial z_j^l}{\partial w_{ij}^l} = a_i^{l-1}, $$ -The derivative of this function with respect to the weights is
- +and
$$ -\frac{\partial{\cal C}(\hat{W^L})}{\partial w_{jk}^L} = \left(a_j^L - t_j\right)\frac{\partial a_j^L}{\partial w_{jk}^{L}}, +\frac{\partial z_j^l}{\partial a_i^{l-1}} = w_{ji}^l. $$ -The last partial derivative can easily be computed and reads (by applying the chain rule)
+With our definition of the activation function we have that (note that this function depends only on \( z_j^l \))
$$ -\frac{\partial a_j^L}{\partial w_{jk}^{L}} = \frac{\partial a_j^L}{\partial z_{j}^{L}}\frac{\partial z_j^L}{\partial w_{jk}^{L}}=a_j^L(1-a_j^L)a_k^{L-1}, +\frac{\partial a_j^l}{\partial z_j^{l}} = a_j^l(1-a_j^l)=f(z_j^l)(1-f(z_j^l)). $$ @@ -357,6 +356,7 @@ $$
-
We have thus
+With these definitions we can now compute the derivative of the cost function in terms of the weights.
+ +Let us specialize to the output layer \( l=L \). Our cost function is
$$ -\frac{\partial{\cal C}(\hat{W^L})}{\partial w_{jk}^L} = \left(a_j^L - t_j\right)a_j^L(1-a_j^L)a_k^{L-1}, +{\cal C}(\hat{W^L}) = \frac{1}{2}\sum_{i=1}^n\left(y_i - t_i\right)^2=\frac{1}{2}\sum_{i=1}^n\left(a_i^L - t_i\right)^2, $$ -Defining
-$$ -\delta_j^L = a_j^L(1-a_j^L)\left(a_j^L - t_j\right) = f'(z_j^L)\frac{\partial {\cal C}}{\partial (a_j^L)}, -$$ - -and using the Hadamard product of two vectors we can write this as
-$$ -\hat{\delta}^L = f'(\hat{z}^L)\circ\frac{\partial {\cal C}}{\partial (\hat{a}^L)}. -$$ - -This is an important expression. The second term on the right handside -measures how fast the cost function is changing as a function of the $j$th -output activation. If, for example, the cost function doesn't depend -much on a particular output node \( j \), then \( \delta_j^L \) will be small, -which is what we would expect. The first term on the right, measures -how fast the activation function \( f \) is changing at a given activation -value \( z_j^L \). -
- -Notice that everything in the above equations is easily computed. In -particular, we compute \( z_j^L \) while computing the behaviour of the -network, and it is only a small additional overhead to compute -\( f'(z^L_j) \). The exact form of the derivative with respect to the -output depends on the form of the cost function. -However, provided the cost function is known there should be little -trouble in calculating -
+The derivative of this function with respect to the weights is
$$ -\frac{\partial {\cal C}}{\partial (a_j^L)} +\frac{\partial{\cal C}(\hat{W^L})}{\partial w_{jk}^L} = \left(a_j^L - t_j\right)\frac{\partial a_j^L}{\partial w_{jk}^{L}}, $$ -With the definition of \( \delta_j^L \) we have a more compact definition of the derivative of the cost function in terms of the weights, namely
+The last partial derivative can easily be computed and reads (by applying the chain rule)
$$ -\frac{\partial{\cal C}(\hat{W^L})}{\partial w_{jk}^L} = \delta_j^La_k^{L-1}. +\frac{\partial a_j^L}{\partial w_{jk}^{L}} = \frac{\partial a_j^L}{\partial z_{j}^{L}}\frac{\partial z_j^L}{\partial w_{jk}^{L}}=a_j^L(1-a_j^L)a_k^{L-1}, $$ @@ -380,6 +358,7 @@ $$
-
It is also easy to see that our previous equation can be written as
+We have thus
$$ -\delta_j^L =\frac{\partial {\cal C}}{\partial z_j^L}= \frac{\partial {\cal C}}{\partial a_j^L}\frac{\partial a_j^L}{\partial z_j^L}, +\frac{\partial{\cal C}(\hat{W^L})}{\partial w_{jk}^L} = \left(a_j^L - t_j\right)a_j^L(1-a_j^L)a_k^{L-1}, $$ -which can also be interpreted as the partial derivative of the cost function with respect to the biases \( b_j^L \), namely
+Defining
$$ -\delta_j^L = \frac{\partial {\cal C}}{\partial b_j^L}\frac{\partial b_j^L}{\partial z_j^L}=\frac{\partial {\cal C}}{\partial b_j^L}, +\delta_j^L = a_j^L(1-a_j^L)\left(a_j^L - t_j\right) = f'(z_j^L)\frac{\partial {\cal C}}{\partial (a_j^L)}, $$ -That is, the error \( \delta_j^L \) is exactly equal to the rate of change of the cost function as a function of the bias.
+and using the Hadamard product of two vectors we can write this as
+$$ +\hat{\delta}^L = f'(\hat{z}^L)\circ\frac{\partial {\cal C}}{\partial (\hat{a}^L)}. +$$ + +This is an important expression. The second term on the right handside +measures how fast the cost function is changing as a function of the $j$th +output activation. If, for example, the cost function doesn't depend +much on a particular output node \( j \), then \( \delta_j^L \) will be small, +which is what we would expect. The first term on the right, measures +how fast the activation function \( f \) is changing at a given activation +value \( z_j^L \). +
+ +Notice that everything in the above equations is easily computed. In +particular, we compute \( z_j^L \) while computing the behaviour of the +network, and it is only a small additional overhead to compute +\( f'(z^L_j) \). The exact form of the derivative with respect to the +output depends on the form of the cost function. +However, provided the cost function is known there should be little +trouble in calculating +
+ +$$ +\frac{\partial {\cal C}}{\partial (a_j^L)} +$$ + +With the definition of \( \delta_j^L \) we have a more compact definition of the derivative of the cost function in terms of the weights, namely
+$$ +\frac{\partial{\cal C}(\hat{W^L})}{\partial w_{jk}^L} = \delta_j^La_k^{L-1}. +$$ + +diff --git a/doc/pub/week40/html/week40-bs.html b/doc/pub/week40/html/week40-bs.html index 76cb74266..dd9e9bdc5 100644 --- a/doc/pub/week40/html/week40-bs.html +++ b/doc/pub/week40/html/week40-bs.html @@ -63,6 +63,7 @@ doconce format html week40.do.txt --html_style=bootstrap --pygments_html_style=d 2, None, 'program-for-stochastic-gradient'), + ('Replace or not', 2, None, 'replace-or-not'), ('Momentum based GD', 2, None, 'momentum-based-gd'), ('More on momentum based approaches', 2, @@ -249,62 +250,63 @@ MathJax.Hub.Config({
In the above code, we have use replacement in setting up the +mini-batches. The discussion +here may be +useful. More material will be added later. +
+In the above code, we have use replacement in setting up the +mini-batches. The discussion +here may be +useful. More material will be added later. +
+In the above code, we have use replacement in setting up the +mini-batches. The discussion +here may be +useful. More material will be added later. +
+