Update on NN

This commit is contained in:
mhjensen
2017-11-22 15:51:08 +01:00
parent 9d2bcb57ab
commit 06d468b2d4
10 changed files with 679 additions and 144 deletions
+127 -26
View File
@@ -47,12 +47,21 @@ Automatically generated HTML file from DocOnce source
('Feed-forward neural networks', 2, None, '___sec4'),
('Recurrent neural networks', 2, None, '___sec5'),
('Other types of networks', 2, None, '___sec6'),
('Mathematical model', 2, None, '___sec7'),
('Mathematical model', 2, None, '___sec8'),
('Multilayer perceptrons', 2, None, '___sec7'),
('Why multilayer perceptrons?', 2, None, '___sec8'),
('Mathematical model', 2, None, '___sec9'),
('Mathematical model', 2, None, '___sec10'),
('Mathematical model', 2, None, '___sec11'),
('Matrix-vector notation', 3, None, '___sec12')]}
('Mathematical model', 2, None, '___sec12'),
('Mathematical model', 2, None, '___sec13'),
('Matrix-vector notation', 3, None, '___sec14'),
('Matrix-vector notation and activation', 3, None, '___sec15'),
('Activation functions', 3, None, '___sec16'),
('Activation functions, Logistic and Hyperbolic ones',
3,
None,
'___sec17'),
('Relevance', 3, None, '___sec18')]}
end of tocinfo -->
<body>
@@ -97,12 +106,18 @@ MathJax.Hub.Config({
<!-- navigation toc: --> <li><a href="#___sec4" style="font-size: 80%;"><b>Feed-forward neural networks</b></a></li>
<!-- navigation toc: --> <li><a href="#___sec5" style="font-size: 80%;"><b>Recurrent neural networks</b></a></li>
<!-- navigation toc: --> <li><a href="#___sec6" style="font-size: 80%;"><b>Other types of networks</b></a></li>
<!-- navigation toc: --> <li><a href="#___sec7" style="font-size: 80%;"><b>Mathematical model</b></a></li>
<!-- navigation toc: --> <li><a href="#___sec8" style="font-size: 80%;"><b>Mathematical model</b></a></li>
<!-- navigation toc: --> <li><a href="#___sec7" style="font-size: 80%;"><b>Multilayer perceptrons</b></a></li>
<!-- navigation toc: --> <li><a href="#___sec8" style="font-size: 80%;"><b>Why multilayer perceptrons?</b></a></li>
<!-- navigation toc: --> <li><a href="#___sec9" style="font-size: 80%;"><b>Mathematical model</b></a></li>
<!-- navigation toc: --> <li><a href="#___sec10" style="font-size: 80%;"><b>Mathematical model</b></a></li>
<!-- navigation toc: --> <li><a href="#___sec11" style="font-size: 80%;"><b>Mathematical model</b></a></li>
<!-- navigation toc: --> <li><a href="#___sec12" style="font-size: 80%;">&nbsp;&nbsp;&nbsp;Matrix-vector notation</a></li>
<!-- navigation toc: --> <li><a href="#___sec12" style="font-size: 80%;"><b>Mathematical model</b></a></li>
<!-- navigation toc: --> <li><a href="#___sec13" style="font-size: 80%;"><b>Mathematical model</b></a></li>
<!-- navigation toc: --> <li><a href="#___sec14" style="font-size: 80%;">&nbsp;&nbsp;&nbsp;Matrix-vector notation</a></li>
<!-- navigation toc: --> <li><a href="#___sec15" style="font-size: 80%;">&nbsp;&nbsp;&nbsp;Matrix-vector notation and activation</a></li>
<!-- navigation toc: --> <li><a href="#___sec16" style="font-size: 80%;">&nbsp;&nbsp;&nbsp;Activation functions</a></li>
<!-- navigation toc: --> <li><a href="#___sec17" style="font-size: 80%;">&nbsp;&nbsp;&nbsp;Activation functions, Logistic and Hyperbolic ones</a></li>
<!-- navigation toc: --> <li><a href="#___sec18" style="font-size: 80%;">&nbsp;&nbsp;&nbsp;Relevance</a></li>
</ul>
</li>
@@ -283,7 +298,41 @@ of how a fully-connected FFNN works, and how it can be used to interpolate data
<p>
<!-- !split -->
<h2 id="___sec7" class="anchor">Mathematical model </h2>
<h2 id="___sec7" class="anchor">Multilayer perceptrons </h2>
<p>
One use often so-called fully-connected feed-forward neural networks with three
or more layers (an input layer, one or more hidden layers and an output layer)
consisting of neurons that have non-linear activation functions.
<p>
Such networks are often called <em>multilayer perceptrons</em> (MLPs)
<p>
<!-- !split -->
<h2 id="___sec8" class="anchor">Why multilayer perceptrons? </h2>
<p>
According to the <em>Universal approximation theorem</em>, a feed-forward neural network with just a single hidden layer containing
a finite number of neurons can approximate a continuous multidimensional function to arbitrary accuracy,
assuming the activation function for the hidden layer is a <b>non-constant, bounded and monotonically-increasing continuous function</b>.
Note that the requirements on the activation function only applies to the hidden layer, the output nodes are always
assumed to be linear, so as to not restrict the range of output values.
<p>
We note that this theorem is only applicable to a NN with <em>one</em> hidden layer.
Therefore, we can easily construct an NN
that employs activation functions which do not satisfy the above requirements, as long as we have at least one layer
with activation functions that <em>do</em>. Furthermore, although the universal approximation theorem
lays the theoretical foundation for regression with neural networks, it does not say anything about how things work in practice:
A neural network can still be able to approximate a given function reasonably well without having the flexibility to fit <em>all other</em>
functions.
<p>
<!-- !split -->
<h2 id="___sec9" class="anchor">Mathematical model </h2>
$$
\begin{equation}
@@ -293,13 +342,13 @@ $$
$$
In an FFNN of such neurons, the <em>inputs</em> \( x_i \)
are the <em>outputs</em> of the neurons in the preceding layer. Furthermore, a MLP is fully-connected,
are the <em>outputs</em> of the neurons in the preceding layer. Furthermore, an MLP is fully-connected,
which means that each neuron receives a weighted sum of the outputs of <em>all</em> neurons in the previous layer.
<p>
<!-- !split -->
<h2 id="___sec8" class="anchor">Mathematical model </h2>
<h2 id="___sec10" class="anchor">Mathematical model </h2>
<p>
First, for each node \( i \) in the first hidden layer, we calculate a weighted sum \( u_i^1 \) of the input coordinates \( x_j \),
@@ -336,7 +385,7 @@ the values of the subsequent layer can be calculated and so forth until the outp
<p>
<!-- !split -->
<h2 id="___sec9" class="anchor">Mathematical model </h2>
<h2 id="___sec11" class="anchor">Mathematical model </h2>
<p>
The output of neuron \( i \) in layer 2 is thus,
@@ -365,7 +414,7 @@ $$
<p>
<!-- !split -->
<h2 id="___sec10" class="anchor">Mathematical model </h2>
<h2 id="___sec12" class="anchor">Mathematical model </h2>
<p>
We can generalize this expression to an MLP with \( l \) hidden layers. The complete functional form
@@ -385,7 +434,7 @@ which illustrates a basic property of MLPs: The only independent variables are t
<p>
<!-- !split -->
<h2 id="___sec11" class="anchor">Mathematical model </h2>
<h2 id="___sec13" class="anchor">Mathematical model </h2>
<p>
This confirms that an MLP,
@@ -407,19 +456,12 @@ $$
where the parameters \( c_i \) are weights and biases. By adjusting these parameters, the activation functions
can be shifted up and down or left and right, change slope or be rescaled
which is the key to the flexibility of a NN.
$$
\begin{equation}
f_o = f(u_o) = u_o
\label{outputActivation}
\end{equation}
$$
which is the key to the flexibility of a neural network.
<p>
<!-- !split -->
<h3 id="___sec12" class="anchor">Matrix-vector notation </h3>
<h3 id="___sec14" class="anchor">Matrix-vector notation </h3>
We can introduce a more convenient notation for the activations in a NN.
@@ -432,7 +474,6 @@ is the bias \( b_i^l \) and activation \( y_i^l \) of node \( i \) in layer \( l
We have that \( \mathrm{W}_l \) is a \( N_{l-1} \times N_l \) matrix, while \( \vec{b}_l \) and \( \vec{y}_l \) are \( N_l \times 1 \) column vectors.
With this notation, the sum in becomes a matrix-vector multiplication, and we can write
the equation for the activations of hidden layer 2 in
$$
\begin{equation}
\vec{y}_2 = f_2(\mathrm{W}_2 \vec{y}_{1} + \vec{b}_{2}) =
@@ -450,27 +491,87 @@ $$
b^2_1 \\
b^2_2 \\
b^2_3 \\
\end{array}\right]\right)
\end{array}\right]\right).
\label{_auto6}
\end{equation}
$$
and we see that the activation of node \( i \) in layer 2 is
<p>
<!-- !split -->
<h3 id="___sec15" class="anchor">Matrix-vector notation and activation </h3>
The activation of node \( i \) in layer 2 is
$$
\begin{equation}
y^2_i = f_2\Bigr(w^2_{i1}y^1_1 + w^2_{i2}y^1_2 + w^2_{i3}y^1_3 + b^2_i\Bigr) =
f_2\left(\sum_{j=1}^3 w^2_{ij} y_j^1 + b^2_i\right)
f_2\left(\sum_{j=1}^3 w^2_{ij} y_j^1 + b^2_i\right).
\label{_auto7}
\end{equation}
$$
which is in accordance with. Note that
This is not just a convenient and compact notation, but also
a useful and intuitive way to think about MLPs: The output is calculated by a series of matrix-vector multiplications
and vector additions that are used as input to the activation functions. For each operation
\( \mathrm{W}_l \vec{y}_{l-1} \) we move forward one layer.
<p>
<!-- !split -->
<h3 id="___sec16" class="anchor">Activation functions </h3>
<p>
A property that characterizes a neural network, other than its connectivity, is the choice of activation function(s).
As described in, the following restrictions are imposed on an activation function for a FFNN
to fulfill the universal approximation theorem
<ul>
<li> Non-constant</li>
<li> Bounded</li>
<li> Monotonically-increasing</li>
<li> Continuous</li>
</ul>
<!-- !split -->
<h3 id="___sec17" class="anchor">Activation functions, Logistic and Hyperbolic ones </h3>
<p>
The second requirement excludes all linear functions. Furthermore, in a MLP with only linear activation functions, each
layer simply performs a linear transformation of its inputs.
<p>
Regardless of the number of layers,
the output of the NN will be nothing but a linear function of the inputs. Thus we need to introduce some kind of
non-linearity to the NN to be able to fit non-linear functions
Typical examples are the logistic <em>Sigmoid</em>
$$
\begin{equation}
f(x) = \frac{1}{1 + e^{-x}},
\label{sigmoidActivationFunction}
\end{equation}
$$
and the <em>hyperbolic tangent</em> function
$$
\begin{equation}
f(x) = \tanh(x)
\label{tanhActivationFunction}
\end{equation}
$$
<p>
<!-- !split -->
<h3 id="___sec18" class="anchor">Relevance </h3>
The <em>sigmoid</em> function are more biologically plausible because
the output of inactive neurons are zero. Such activation function are called <em>one-sided</em>. However,
it has been shown that the hyperbolic tangent
performs better than the sigmoid for training MLPs.
has become the most popular for <em>deep neural networks</em>
<p>
<!-- ------------------- end of main content --------------- -->
+124 -27
View File
@@ -304,7 +304,41 @@ of how a fully-connected FFNN works, and how it can be used to interpolate data
<section>
<h2 id="___sec7">Mathematical model </h2>
<h2 id="___sec7">Multilayer perceptrons </h2>
<p>
One use often so-called fully-connected feed-forward neural networks with three
or more layers (an input layer, one or more hidden layers and an output layer)
consisting of neurons that have non-linear activation functions.
<p>
Such networks are often called <em>multilayer perceptrons</em> (MLPs)
</section>
<section>
<h2 id="___sec8">Why multilayer perceptrons? </h2>
<p>
According to the <em>Universal approximation theorem</em>, a feed-forward neural network with just a single hidden layer containing
a finite number of neurons can approximate a continuous multidimensional function to arbitrary accuracy,
assuming the activation function for the hidden layer is a <b>non-constant, bounded and monotonically-increasing continuous function</b>.
Note that the requirements on the activation function only applies to the hidden layer, the output nodes are always
assumed to be linear, so as to not restrict the range of output values.
<p>
We note that this theorem is only applicable to a NN with <em>one</em> hidden layer.
Therefore, we can easily construct an NN
that employs activation functions which do not satisfy the above requirements, as long as we have at least one layer
with activation functions that <em>do</em>. Furthermore, although the universal approximation theorem
lays the theoretical foundation for regression with neural networks, it does not say anything about how things work in practice:
A neural network can still be able to approximate a given function reasonably well without having the flexibility to fit <em>all other</em>
functions.
</section>
<section>
<h2 id="___sec9">Mathematical model </h2>
<p>&nbsp;<br>
$$
@@ -316,13 +350,13 @@ $$
<p>&nbsp;<br>
In an FFNN of such neurons, the <em>inputs</em> \( x_i \)
are the <em>outputs</em> of the neurons in the preceding layer. Furthermore, a MLP is fully-connected,
are the <em>outputs</em> of the neurons in the preceding layer. Furthermore, an MLP is fully-connected,
which means that each neuron receives a weighted sum of the outputs of <em>all</em> neurons in the previous layer.
</section>
<section>
<h2 id="___sec8">Mathematical model </h2>
<h2 id="___sec10">Mathematical model </h2>
<p>
First, for each node \( i \) in the first hidden layer, we calculate a weighted sum \( u_i^1 \) of the input coordinates \( x_j \),
@@ -365,7 +399,7 @@ the values of the subsequent layer can be calculated and so forth until the outp
<section>
<h2 id="___sec9">Mathematical model </h2>
<h2 id="___sec11">Mathematical model </h2>
<p>
The output of neuron \( i \) in layer 2 is thus,
@@ -398,7 +432,7 @@ $$
<section>
<h2 id="___sec10">Mathematical model </h2>
<h2 id="___sec12">Mathematical model </h2>
<p>
We can generalize this expression to an MLP with \( l \) hidden layers. The complete functional form
@@ -420,7 +454,7 @@ which illustrates a basic property of MLPs: The only independent variables are t
<section>
<h2 id="___sec11">Mathematical model </h2>
<h2 id="___sec13">Mathematical model </h2>
<p>
This confirms that an MLP,
@@ -444,22 +478,13 @@ $$
where the parameters \( c_i \) are weights and biases. By adjusting these parameters, the activation functions
can be shifted up and down or left and right, change slope or be rescaled
which is the key to the flexibility of a NN.
<p>&nbsp;<br>
$$
\begin{equation}
f_o = f(u_o) = u_o
\tag{12}
\end{equation}
$$
<p>&nbsp;<br>
which is the key to the flexibility of a neural network.
</section>
<section>
<h3 id="___sec12">Matrix-vector notation </h3>
<h3 id="___sec14">Matrix-vector notation </h3>
We can introduce a more convenient notation for the activations in a NN.
@@ -472,7 +497,6 @@ is the bias \( b_i^l \) and activation \( y_i^l \) of node \( i \) in layer \( l
We have that \( \mathrm{W}_l \) is a \( N_{l-1} \times N_l \) matrix, while \( \vec{b}_l \) and \( \vec{y}_l \) are \( N_l \times 1 \) column vectors.
With this notation, the sum in becomes a matrix-vector multiplication, and we can write
the equation for the activations of hidden layer 2 in
<p>&nbsp;<br>
$$
\begin{equation}
@@ -491,29 +515,102 @@ $$
b^2_1 \\
b^2_2 \\
b^2_3 \\
\end{array}\right]\right)
\tag{13}
\end{array}\right]\right).
\tag{12}
\end{equation}
$$
<p>&nbsp;<br>
</section>
and we see that the activation of node \( i \) in layer 2 is
<section>
<h3 id="___sec15">Matrix-vector notation and activation </h3>
The activation of node \( i \) in layer 2 is
<p>&nbsp;<br>
$$
\begin{equation}
y^2_i = f_2\Bigr(w^2_{i1}y^1_1 + w^2_{i2}y^1_2 + w^2_{i3}y^1_3 + b^2_i\Bigr) =
f_2\left(\sum_{j=1}^3 w^2_{ij} y_j^1 + b^2_i\right)
f_2\left(\sum_{j=1}^3 w^2_{ij} y_j^1 + b^2_i\right).
\tag{13}
\end{equation}
$$
<p>&nbsp;<br>
This is not just a convenient and compact notation, but also
a useful and intuitive way to think about MLPs: The output is calculated by a series of matrix-vector multiplications
and vector additions that are used as input to the activation functions. For each operation
\( \mathrm{W}_l \vec{y}_{l-1} \) we move forward one layer.
</section>
<section>
<h3 id="___sec16">Activation functions </h3>
<p>
A property that characterizes a neural network, other than its connectivity, is the choice of activation function(s).
As described in, the following restrictions are imposed on an activation function for a FFNN
to fulfill the universal approximation theorem
<ul>
<p><li> Non-constant</li>
<p><li> Bounded</li>
<p><li> Monotonically-increasing</li>
<p><li> Continuous</li>
</ul>
</section>
<section>
<h3 id="___sec17">Activation functions, Logistic and Hyperbolic ones </h3>
<p>
The second requirement excludes all linear functions. Furthermore, in a MLP with only linear activation functions, each
layer simply performs a linear transformation of its inputs.
<p>
Regardless of the number of layers,
the output of the NN will be nothing but a linear function of the inputs. Thus we need to introduce some kind of
non-linearity to the NN to be able to fit non-linear functions
Typical examples are the logistic <em>Sigmoid</em>
<p>&nbsp;<br>
$$
\begin{equation}
f(x) = \frac{1}{1 + e^{-x}},
\tag{14}
\end{equation}
$$
<p>&nbsp;<br>
which is in accordance with. Note that
This is not just a convenient and compact notation, but also
a useful and intuitive way to think about MLPs: The output is calculated by a series of matrix-vector multiplications
and vector additions that are used as input to the activation functions. For each operation
\( \mathrm{W}_l \vec{y}_{l-1} \) we move forward one layer.
and the <em>hyperbolic tangent</em> function
<p>&nbsp;<br>
$$
\begin{equation}
f(x) = \tanh(x)
\tag{15}
\end{equation}
$$
<p>&nbsp;<br>
</section>
<section>
<h3 id="___sec18">Relevance </h3>
The <em>sigmoid</em> function are more biologically plausible because
the output of inactive neurons are zero. Such activation function are called <em>one-sided</em>. However,
it has been shown that the hyperbolic tangent
performs better than the sigmoid for training MLPs.
has become the most popular for <em>deep neural networks</em>
</section>
+118 -23
View File
@@ -41,12 +41,21 @@ div { text-align: justify; text-justify: inter-word; }
('Feed-forward neural networks', 2, None, '___sec4'),
('Recurrent neural networks', 2, None, '___sec5'),
('Other types of networks', 2, None, '___sec6'),
('Mathematical model', 2, None, '___sec7'),
('Mathematical model', 2, None, '___sec8'),
('Multilayer perceptrons', 2, None, '___sec7'),
('Why multilayer perceptrons?', 2, None, '___sec8'),
('Mathematical model', 2, None, '___sec9'),
('Mathematical model', 2, None, '___sec10'),
('Mathematical model', 2, None, '___sec11'),
('Matrix-vector notation', 3, None, '___sec12')]}
('Mathematical model', 2, None, '___sec12'),
('Mathematical model', 2, None, '___sec13'),
('Matrix-vector notation', 3, None, '___sec14'),
('Matrix-vector notation and activation', 3, None, '___sec15'),
('Activation functions', 3, None, '___sec16'),
('Activation functions, Logistic and Hyperbolic ones',
3,
None,
'___sec17'),
('Relevance', 3, None, '___sec18')]}
end of tocinfo -->
<body>
@@ -232,7 +241,41 @@ of how a fully-connected FFNN works, and how it can be used to interpolate data
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec7">Mathematical model </h2>
<h2 id="___sec7">Multilayer perceptrons </h2>
<p>
One use often so-called fully-connected feed-forward neural networks with three
or more layers (an input layer, one or more hidden layers and an output layer)
consisting of neurons that have non-linear activation functions.
<p>
Such networks are often called <em>multilayer perceptrons</em> (MLPs)
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec8">Why multilayer perceptrons? </h2>
<p>
According to the <em>Universal approximation theorem</em>, a feed-forward neural network with just a single hidden layer containing
a finite number of neurons can approximate a continuous multidimensional function to arbitrary accuracy,
assuming the activation function for the hidden layer is a <b>non-constant, bounded and monotonically-increasing continuous function</b>.
Note that the requirements on the activation function only applies to the hidden layer, the output nodes are always
assumed to be linear, so as to not restrict the range of output values.
<p>
We note that this theorem is only applicable to a NN with <em>one</em> hidden layer.
Therefore, we can easily construct an NN
that employs activation functions which do not satisfy the above requirements, as long as we have at least one layer
with activation functions that <em>do</em>. Furthermore, although the universal approximation theorem
lays the theoretical foundation for regression with neural networks, it does not say anything about how things work in practice:
A neural network can still be able to approximate a given function reasonably well without having the flexibility to fit <em>all other</em>
functions.
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec9">Mathematical model </h2>
$$
\begin{equation}
@@ -242,13 +285,13 @@ $$
$$
In an FFNN of such neurons, the <em>inputs</em> \( x_i \)
are the <em>outputs</em> of the neurons in the preceding layer. Furthermore, a MLP is fully-connected,
are the <em>outputs</em> of the neurons in the preceding layer. Furthermore, an MLP is fully-connected,
which means that each neuron receives a weighted sum of the outputs of <em>all</em> neurons in the previous layer.
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec8">Mathematical model </h2>
<h2 id="___sec10">Mathematical model </h2>
<p>
First, for each node \( i \) in the first hidden layer, we calculate a weighted sum \( u_i^1 \) of the input coordinates \( x_j \),
@@ -285,7 +328,7 @@ the values of the subsequent layer can be calculated and so forth until the outp
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec9">Mathematical model </h2>
<h2 id="___sec11">Mathematical model </h2>
<p>
The output of neuron \( i \) in layer 2 is thus,
@@ -314,7 +357,7 @@ $$
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec10">Mathematical model </h2>
<h2 id="___sec12">Mathematical model </h2>
<p>
We can generalize this expression to an MLP with \( l \) hidden layers. The complete functional form
@@ -334,7 +377,7 @@ which illustrates a basic property of MLPs: The only independent variables are t
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec11">Mathematical model </h2>
<h2 id="___sec13">Mathematical model </h2>
<p>
This confirms that an MLP,
@@ -356,19 +399,12 @@ $$
where the parameters \( c_i \) are weights and biases. By adjusting these parameters, the activation functions
can be shifted up and down or left and right, change slope or be rescaled
which is the key to the flexibility of a NN.
$$
\begin{equation}
f_o = f(u_o) = u_o
\label{outputActivation}
\end{equation}
$$
which is the key to the flexibility of a neural network.
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h3 id="___sec12">Matrix-vector notation </h3>
<h3 id="___sec14">Matrix-vector notation </h3>
We can introduce a more convenient notation for the activations in a NN.
@@ -381,7 +417,6 @@ is the bias \( b_i^l \) and activation \( y_i^l \) of node \( i \) in layer \( l
We have that \( \mathrm{W}_l \) is a \( N_{l-1} \times N_l \) matrix, while \( \vec{b}_l \) and \( \vec{y}_l \) are \( N_l \times 1 \) column vectors.
With this notation, the sum in becomes a matrix-vector multiplication, and we can write
the equation for the activations of hidden layer 2 in
$$
\begin{equation}
\vec{y}_2 = f_2(\mathrm{W}_2 \vec{y}_{1} + \vec{b}_{2}) =
@@ -399,27 +434,87 @@ $$
b^2_1 \\
b^2_2 \\
b^2_3 \\
\end{array}\right]\right)
\end{array}\right]\right).
\label{_auto6}
\end{equation}
$$
and we see that the activation of node \( i \) in layer 2 is
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h3 id="___sec15">Matrix-vector notation and activation </h3>
The activation of node \( i \) in layer 2 is
$$
\begin{equation}
y^2_i = f_2\Bigr(w^2_{i1}y^1_1 + w^2_{i2}y^1_2 + w^2_{i3}y^1_3 + b^2_i\Bigr) =
f_2\left(\sum_{j=1}^3 w^2_{ij} y_j^1 + b^2_i\right)
f_2\left(\sum_{j=1}^3 w^2_{ij} y_j^1 + b^2_i\right).
\label{_auto7}
\end{equation}
$$
which is in accordance with. Note that
This is not just a convenient and compact notation, but also
a useful and intuitive way to think about MLPs: The output is calculated by a series of matrix-vector multiplications
and vector additions that are used as input to the activation functions. For each operation
\( \mathrm{W}_l \vec{y}_{l-1} \) we move forward one layer.
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h3 id="___sec16">Activation functions </h3>
<p>
A property that characterizes a neural network, other than its connectivity, is the choice of activation function(s).
As described in, the following restrictions are imposed on an activation function for a FFNN
to fulfill the universal approximation theorem
<ul>
<li> Non-constant</li>
<li> Bounded</li>
<li> Monotonically-increasing</li>
<li> Continuous</li>
</ul>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h3 id="___sec17">Activation functions, Logistic and Hyperbolic ones </h3>
<p>
The second requirement excludes all linear functions. Furthermore, in a MLP with only linear activation functions, each
layer simply performs a linear transformation of its inputs.
<p>
Regardless of the number of layers,
the output of the NN will be nothing but a linear function of the inputs. Thus we need to introduce some kind of
non-linearity to the NN to be able to fit non-linear functions
Typical examples are the logistic <em>Sigmoid</em>
$$
\begin{equation}
f(x) = \frac{1}{1 + e^{-x}},
\label{sigmoidActivationFunction}
\end{equation}
$$
and the <em>hyperbolic tangent</em> function
$$
\begin{equation}
f(x) = \tanh(x)
\label{tanhActivationFunction}
\end{equation}
$$
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h3 id="___sec18">Relevance </h3>
The <em>sigmoid</em> function are more biologically plausible because
the output of inactive neurons are zero. Such activation function are called <em>one-sided</em>. However,
it has been shown that the hyperbolic tangent
performs better than the sigmoid for training MLPs.
has become the most popular for <em>deep neural networks</em>
<p>
<!-- ------------------- end of main content --------------- -->
+118 -23
View File
@@ -46,12 +46,21 @@ div { text-align: justify; text-justify: inter-word; }
('Feed-forward neural networks', 2, None, '___sec4'),
('Recurrent neural networks', 2, None, '___sec5'),
('Other types of networks', 2, None, '___sec6'),
('Mathematical model', 2, None, '___sec7'),
('Mathematical model', 2, None, '___sec8'),
('Multilayer perceptrons', 2, None, '___sec7'),
('Why multilayer perceptrons?', 2, None, '___sec8'),
('Mathematical model', 2, None, '___sec9'),
('Mathematical model', 2, None, '___sec10'),
('Mathematical model', 2, None, '___sec11'),
('Matrix-vector notation', 3, None, '___sec12')]}
('Mathematical model', 2, None, '___sec12'),
('Mathematical model', 2, None, '___sec13'),
('Matrix-vector notation', 3, None, '___sec14'),
('Matrix-vector notation and activation', 3, None, '___sec15'),
('Activation functions', 3, None, '___sec16'),
('Activation functions, Logistic and Hyperbolic ones',
3,
None,
'___sec17'),
('Relevance', 3, None, '___sec18')]}
end of tocinfo -->
<body>
@@ -237,7 +246,41 @@ of how a fully-connected FFNN works, and how it can be used to interpolate data
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec7">Mathematical model </h2>
<h2 id="___sec7">Multilayer perceptrons </h2>
<p>
One use often so-called fully-connected feed-forward neural networks with three
or more layers (an input layer, one or more hidden layers and an output layer)
consisting of neurons that have non-linear activation functions.
<p>
Such networks are often called <em>multilayer perceptrons</em> (MLPs)
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec8">Why multilayer perceptrons? </h2>
<p>
According to the <em>Universal approximation theorem</em>, a feed-forward neural network with just a single hidden layer containing
a finite number of neurons can approximate a continuous multidimensional function to arbitrary accuracy,
assuming the activation function for the hidden layer is a <b>non-constant, bounded and monotonically-increasing continuous function</b>.
Note that the requirements on the activation function only applies to the hidden layer, the output nodes are always
assumed to be linear, so as to not restrict the range of output values.
<p>
We note that this theorem is only applicable to a NN with <em>one</em> hidden layer.
Therefore, we can easily construct an NN
that employs activation functions which do not satisfy the above requirements, as long as we have at least one layer
with activation functions that <em>do</em>. Furthermore, although the universal approximation theorem
lays the theoretical foundation for regression with neural networks, it does not say anything about how things work in practice:
A neural network can still be able to approximate a given function reasonably well without having the flexibility to fit <em>all other</em>
functions.
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec9">Mathematical model </h2>
$$
\begin{equation}
@@ -247,13 +290,13 @@ $$
$$
In an FFNN of such neurons, the <em>inputs</em> \( x_i \)
are the <em>outputs</em> of the neurons in the preceding layer. Furthermore, a MLP is fully-connected,
are the <em>outputs</em> of the neurons in the preceding layer. Furthermore, an MLP is fully-connected,
which means that each neuron receives a weighted sum of the outputs of <em>all</em> neurons in the previous layer.
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec8">Mathematical model </h2>
<h2 id="___sec10">Mathematical model </h2>
<p>
First, for each node \( i \) in the first hidden layer, we calculate a weighted sum \( u_i^1 \) of the input coordinates \( x_j \),
@@ -290,7 +333,7 @@ the values of the subsequent layer can be calculated and so forth until the outp
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec9">Mathematical model </h2>
<h2 id="___sec11">Mathematical model </h2>
<p>
The output of neuron \( i \) in layer 2 is thus,
@@ -319,7 +362,7 @@ $$
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec10">Mathematical model </h2>
<h2 id="___sec12">Mathematical model </h2>
<p>
We can generalize this expression to an MLP with \( l \) hidden layers. The complete functional form
@@ -339,7 +382,7 @@ which illustrates a basic property of MLPs: The only independent variables are t
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h2 id="___sec11">Mathematical model </h2>
<h2 id="___sec13">Mathematical model </h2>
<p>
This confirms that an MLP,
@@ -361,19 +404,12 @@ $$
where the parameters \( c_i \) are weights and biases. By adjusting these parameters, the activation functions
can be shifted up and down or left and right, change slope or be rescaled
which is the key to the flexibility of a NN.
$$
\begin{equation}
f_o = f(u_o) = u_o
\label{outputActivation}
\end{equation}
$$
which is the key to the flexibility of a neural network.
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h3 id="___sec12">Matrix-vector notation </h3>
<h3 id="___sec14">Matrix-vector notation </h3>
We can introduce a more convenient notation for the activations in a NN.
@@ -386,7 +422,6 @@ is the bias \( b_i^l \) and activation \( y_i^l \) of node \( i \) in layer \( l
We have that \( \mathrm{W}_l \) is a \( N_{l-1} \times N_l \) matrix, while \( \vec{b}_l \) and \( \vec{y}_l \) are \( N_l \times 1 \) column vectors.
With this notation, the sum in becomes a matrix-vector multiplication, and we can write
the equation for the activations of hidden layer 2 in
$$
\begin{equation}
\vec{y}_2 = f_2(\mathrm{W}_2 \vec{y}_{1} + \vec{b}_{2}) =
@@ -404,27 +439,87 @@ $$
b^2_1 \\
b^2_2 \\
b^2_3 \\
\end{array}\right]\right)
\end{array}\right]\right).
\label{_auto6}
\end{equation}
$$
and we see that the activation of node \( i \) in layer 2 is
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h3 id="___sec15">Matrix-vector notation and activation </h3>
The activation of node \( i \) in layer 2 is
$$
\begin{equation}
y^2_i = f_2\Bigr(w^2_{i1}y^1_1 + w^2_{i2}y^1_2 + w^2_{i3}y^1_3 + b^2_i\Bigr) =
f_2\left(\sum_{j=1}^3 w^2_{ij} y_j^1 + b^2_i\right)
f_2\left(\sum_{j=1}^3 w^2_{ij} y_j^1 + b^2_i\right).
\label{_auto7}
\end{equation}
$$
which is in accordance with. Note that
This is not just a convenient and compact notation, but also
a useful and intuitive way to think about MLPs: The output is calculated by a series of matrix-vector multiplications
and vector additions that are used as input to the activation functions. For each operation
\( \mathrm{W}_l \vec{y}_{l-1} \) we move forward one layer.
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h3 id="___sec16">Activation functions </h3>
<p>
A property that characterizes a neural network, other than its connectivity, is the choice of activation function(s).
As described in, the following restrictions are imposed on an activation function for a FFNN
to fulfill the universal approximation theorem
<ul>
<li> Non-constant</li>
<li> Bounded</li>
<li> Monotonically-increasing</li>
<li> Continuous</li>
</ul>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h3 id="___sec17">Activation functions, Logistic and Hyperbolic ones </h3>
<p>
The second requirement excludes all linear functions. Furthermore, in a MLP with only linear activation functions, each
layer simply performs a linear transformation of its inputs.
<p>
Regardless of the number of layers,
the output of the NN will be nothing but a linear function of the inputs. Thus we need to introduce some kind of
non-linearity to the NN to be able to fit non-linear functions
Typical examples are the logistic <em>Sigmoid</em>
$$
\begin{equation}
f(x) = \frac{1}{1 + e^{-x}},
\label{sigmoidActivationFunction}
\end{equation}
$$
and the <em>hyperbolic tangent</em> function
$$
\begin{equation}
f(x) = \tanh(x)
\label{tanhActivationFunction}
\end{equation}
$$
<p>
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
<h3 id="___sec18">Relevance </h3>
The <em>sigmoid</em> function are more biologically plausible because
the output of inactive neurons are zero. Such activation function are called <em>one-sided</em>. However,
it has been shown that the hyperbolic tangent
performs better than the sigmoid for training MLPs.
has become the most popular for <em>deep neural networks</em>
<p>
<!-- ------------------- end of main content --------------- -->
+109 -29
View File
@@ -142,6 +142,30 @@
"Other types of NNs could also be mentioned, but are outside the scope of this work. We will now move on to a detailed description\n",
"of how a fully-connected FFNN works, and how it can be used to interpolate data sets. \n",
"\n",
"## Multilayer perceptrons\n",
"\n",
"One use often so-called fully-connected feed-forward neural networks with three\n",
"or more layers (an input layer, one or more hidden layers and an output layer)\n",
"consisting of neurons that have non-linear activation functions.\n",
"\n",
"Such networks are often called *multilayer perceptrons* (MLPs)\n",
"\n",
"## Why multilayer perceptrons?\n",
"\n",
"According to the *Universal approximation theorem*, a feed-forward neural network with just a single hidden layer containing \n",
"a finite number of neurons can approximate a continuous multidimensional function to arbitrary accuracy, \n",
"assuming the activation function for the hidden layer is a **non-constant, bounded and monotonically-increasing continuous function**.\n",
"Note that the requirements on the activation function only applies to the hidden layer, the output nodes are always\n",
"assumed to be linear, so as to not restrict the range of output values. \n",
"\n",
"We note that this theorem is only applicable to a NN with *one* hidden layer. \n",
"Therefore, we can easily construct an NN \n",
"that employs activation functions which do not satisfy the above requirements, as long as we have at least one layer\n",
"with activation functions that *do*. Furthermore, although the universal approximation theorem\n",
"lays the theoretical foundation for regression with neural networks, it does not say anything about how things work in practice: \n",
"A neural network can still be able to approximate a given function reasonably well without having the flexibility to fit *all other*\n",
"functions. \n",
"\n",
"\n",
"\n",
"## Mathematical model"
@@ -167,7 +191,7 @@
"metadata": {},
"source": [
"In an FFNN of such neurons, the *inputs* $x_i$\n",
"are the *outputs* of the neurons in the preceding layer. Furthermore, a MLP is fully-connected, \n",
"are the *outputs* of the neurons in the preceding layer. Furthermore, an MLP is fully-connected, \n",
"which means that each neuron receives a weighted sum of the outputs of *all* neurons in the previous layer. \n",
"\n",
"## Mathematical model\n",
@@ -384,28 +408,8 @@
"source": [
"where the parameters $c_i$ are weights and biases. By adjusting these parameters, the activation functions\n",
"can be shifted up and down or left and right, change slope or be rescaled \n",
"which is the key to the flexibility of a NN."
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"<!-- Equation labels as ordinary links -->\n",
"<div id=\"outputActivation\"></div>\n",
"which is the key to the flexibility of a neural network. \n",
"\n",
"$$\n",
"\\begin{equation}\n",
" f_o = f(u_o) = u_o\n",
"\\label{outputActivation} \\tag{12}\n",
"\\end{equation}\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Matrix-vector notation\n",
"\n",
"We can introduce a more convenient notation for the activations in a NN. \n",
@@ -443,8 +447,8 @@
" b^2_1 \\\\\n",
" b^2_2 \\\\\n",
" b^2_3 \\\\\n",
" \\end{array}\\right]\\right)\n",
"\\label{_auto6} \\tag{13}\n",
" \\end{array}\\right]\\right).\n",
"\\label{_auto6} \\tag{12}\n",
"\\end{equation}\n",
"$$"
]
@@ -453,7 +457,9 @@
"cell_type": "markdown",
"metadata": {},
"source": [
"and we see that the activation of node $i$ in layer 2 is"
"### Matrix-vector notation and activation\n",
"\n",
"The activation of node $i$ in layer 2 is"
]
},
{
@@ -466,8 +472,8 @@
"$$\n",
"\\begin{equation}\n",
" y^2_i = f_2\\Bigr(w^2_{i1}y^1_1 + w^2_{i2}y^1_2 + w^2_{i3}y^1_3 + b^2_i\\Bigr) = \n",
" f_2\\left(\\sum_{j=1}^3 w^2_{ij} y_j^1 + b^2_i\\right)\n",
"\\label{_auto7} \\tag{14}\n",
" f_2\\left(\\sum_{j=1}^3 w^2_{ij} y_j^1 + b^2_i\\right).\n",
"\\label{_auto7} \\tag{13}\n",
"\\end{equation}\n",
"$$"
]
@@ -476,11 +482,85 @@
"cell_type": "markdown",
"metadata": {},
"source": [
"which is in accordance with. Note that\n",
"This is not just a convenient and compact notation, but also \n",
"a useful and intuitive way to think about MLPs: The output is calculated by a series of matrix-vector multiplications\n",
"and vector additions that are used as input to the activation functions. For each operation \n",
"$\\mathrm{W}_l \\vec{y}_{l-1}$ we move forward one layer."
"$\\mathrm{W}_l \\vec{y}_{l-1}$ we move forward one layer. \n",
"\n",
"\n",
"### Activation functions\n",
"\n",
"A property that characterizes a neural network, other than its connectivity, is the choice of activation function(s). \n",
"As described in, the following restrictions are imposed on an activation function for a FFNN\n",
"to fulfill the universal approximation theorem\n",
"\n",
" * Non-constant\n",
"\n",
" * Bounded\n",
"\n",
" * Monotonically-increasing\n",
"\n",
" * Continuous\n",
"\n",
"### Activation functions, Logistic and Hyperbolic ones\n",
"\n",
"The second requirement excludes all linear functions. Furthermore, in a MLP with only linear activation functions, each \n",
"layer simply performs a linear transformation of its inputs.\n",
"\n",
"Regardless of the number of layers, \n",
"the output of the NN will be nothing but a linear function of the inputs. Thus we need to introduce some kind of \n",
"non-linearity to the NN to be able to fit non-linear functions\n",
"Typical examples are the logistic *Sigmoid*"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"<!-- Equation labels as ordinary links -->\n",
"<div id=\"sigmoidActivationFunction\"></div>\n",
"\n",
"$$\n",
"\\begin{equation}\n",
" f(x) = \\frac{1}{1 + e^{-x}},\n",
"\\label{sigmoidActivationFunction} \\tag{14}\n",
"\\end{equation}\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"and the *hyperbolic tangent* function"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"<!-- Equation labels as ordinary links -->\n",
"<div id=\"tanhActivationFunction\"></div>\n",
"\n",
"$$\n",
"\\begin{equation}\n",
" f(x) = \\tanh(x)\n",
"\\label{tanhActivationFunction} \\tag{15}\n",
"\\end{equation}\n",
"$$"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Relevance\n",
"\n",
"The *sigmoid* function are more biologically plausible because \n",
"the output of inactive neurons are zero. Such activation function are called *one-sided*. However,\n",
"it has been shown that the hyperbolic tangent \n",
"performs better than the sigmoid for training MLPs. \n",
"has become the most popular for *deep neural networks*"
]
}
],
Binary file not shown.
Binary file not shown.
Binary file not shown.
+83 -16
View File
@@ -122,6 +122,32 @@ a separate type of NN due the unusual activation functions.
Other types of NNs could also be mentioned, but are outside the scope of this work. We will now move on to a detailed description
of how a fully-connected FFNN works, and how it can be used to interpolate data sets.
!split
===== Multilayer perceptrons =====
One use often so-called fully-connected feed-forward neural networks with three
or more layers (an input layer, one or more hidden layers and an output layer)
consisting of neurons that have non-linear activation functions.
Such networks are often called *multilayer perceptrons* (MLPs)
!split
===== Why multilayer perceptrons? =====
According to the *Universal approximation theorem*, a feed-forward neural network with just a single hidden layer containing
a finite number of neurons can approximate a continuous multidimensional function to arbitrary accuracy,
assuming the activation function for the hidden layer is a _non-constant, bounded and monotonically-increasing continuous function_.
Note that the requirements on the activation function only applies to the hidden layer, the output nodes are always
assumed to be linear, so as to not restrict the range of output values.
We note that this theorem is only applicable to a NN with *one* hidden layer.
Therefore, we can easily construct an NN
that employs activation functions which do not satisfy the above requirements, as long as we have at least one layer
with activation functions that *do*. Furthermore, although the universal approximation theorem
lays the theoretical foundation for regression with neural networks, it does not say anything about how things work in practice:
A neural network can still be able to approximate a given function reasonably well without having the flexibility to fit *all other*
functions.
!split
@@ -134,7 +160,7 @@ of how a fully-connected FFNN works, and how it can be used to interpolate data
\end{equation}
!et
In an FFNN of such neurons, the *inputs* $x_i$
are the *outputs* of the neurons in the preceding layer. Furthermore, a MLP is fully-connected,
are the *outputs* of the neurons in the preceding layer. Furthermore, an MLP is fully-connected,
which means that each neuron receives a weighted sum of the outputs of *all* neurons in the previous layer.
!split
@@ -226,16 +252,7 @@ the expression is essentially a nested sum of scaled activation functions of th
!et
where the parameters $c_i$ are weights and biases. By adjusting these parameters, the activation functions
can be shifted up and down or left and right, change slope or be rescaled
which is the key to the flexibility of a NN.
!bt
\begin{equation}
f_o = f(u_o) = u_o
label{outputActivation}
\end{equation}
!et
which is the key to the flexibility of a neural network.
!split
=== Matrix-vector notation ===
@@ -248,7 +265,6 @@ is the bias $b_i^l$ and activation $y_i^l$ of node $i$ in layer $l$ respectively
We have that $\mathrm{W}_l$ is a $N_{l-1} \times N_l$ matrix, while $\vec{b}_l$ and $\vec{y}_l$ are $N_l \times 1$ column vectors.
With this notation, the sum in becomes a matrix-vector multiplication, and we can write
the equation for the activations of hidden layer 2 in
!bt
\begin{equation}
\vec{y}_2 = f_2(\mathrm{W}_2 \vec{y}_{1} + \vec{b}_{2}) =
@@ -266,21 +282,72 @@ the equation for the activations of hidden layer 2 in
b^2_1 \\
b^2_2 \\
b^2_3 \\
\end{array}\right]\right)
\end{array}\right]\right).
\end{equation}
!et
and we see that the activation of node $i$ in layer 2 is
!split
=== Matrix-vector notation and activation ===
The activation of node $i$ in layer 2 is
!bt
\begin{equation}
y^2_i = f_2\Bigr(w^2_{i1}y^1_1 + w^2_{i2}y^1_2 + w^2_{i3}y^1_3 + b^2_i\Bigr) =
f_2\left(\sum_{j=1}^3 w^2_{ij} y_j^1 + b^2_i\right)
f_2\left(\sum_{j=1}^3 w^2_{ij} y_j^1 + b^2_i\right).
\end{equation}
!et
which is in accordance with. Note that
This is not just a convenient and compact notation, but also
a useful and intuitive way to think about MLPs: The output is calculated by a series of matrix-vector multiplications
and vector additions that are used as input to the activation functions. For each operation
$\mathrm{W}_l \vec{y}_{l-1}$ we move forward one layer.
!split
=== Activation functions ===
A property that characterizes a neural network, other than its connectivity, is the choice of activation function(s).
As described in, the following restrictions are imposed on an activation function for a FFNN
to fulfill the universal approximation theorem
* Non-constant
* Bounded
* Monotonically-increasing
* Continuous
!split
=== Activation functions, Logistic and Hyperbolic ones ===
The second requirement excludes all linear functions. Furthermore, in a MLP with only linear activation functions, each
layer simply performs a linear transformation of its inputs.
Regardless of the number of layers,
the output of the NN will be nothing but a linear function of the inputs. Thus we need to introduce some kind of
non-linearity to the NN to be able to fit non-linear functions
Typical examples are the logistic *Sigmoid*
!bt
\begin{equation}
f(x) = \frac{1}{1 + e^{-x}},
label{sigmoidActivationFunction}
\end{equation}
!et
and the *hyperbolic tangent* function
!bt
\begin{equation}
f(x) = \tanh(x)
label{tanhActivationFunction}
\end{equation}
!et
!split
=== Relevance ===
The *sigmoid* function are more biologically plausible because
the output of inactive neurons are zero. Such activation function are called *one-sided*. However,
it has been shown that the hyperbolic tangent
performs better than the sigmoid for training MLPs.
has become the most popular for *deep neural networks*