cleaning up
This commit is contained in:
@@ -269,19 +269,16 @@ means that each node in the output layer has a linear activation function). The
|
||||
there are no cycles, thus RBFs can be viewed as a type of fully-connected FFNN. They are however usually treated as
|
||||
a separate type of NN due the unusual activation functions.
|
||||
|
||||
<p>
|
||||
Other types of NNs could also be mentioned, but are outside the scope of this work. We will now move on to a detailed description
|
||||
of how a fully-connected FFNN works, and how it can be used to interpolate data sets.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec6">Multilayer perceptrons </h2>
|
||||
|
||||
<p>
|
||||
One use often so-called fully-connected feed-forward neural networks with three
|
||||
or more layers (an input layer, one or more hidden layers and an output layer)
|
||||
consisting of neurons that have non-linear activation functions.
|
||||
One uses often so-called fully-connected feed-forward neural networks
|
||||
with three or more layers (an input layer, one or more hidden layers
|
||||
and an output layer) consisting of neurons that have non-linear
|
||||
activation functions.
|
||||
|
||||
<p>
|
||||
Such networks are often called <em>multilayer perceptrons</em> (MLPs)
|
||||
@@ -295,17 +292,22 @@ Such networks are often called <em>multilayer perceptrons</em> (MLPs)
|
||||
According to the <em>Universal approximation theorem</em>, a feed-forward neural network with just a single hidden layer containing
|
||||
a finite number of neurons can approximate a continuous multidimensional function to arbitrary accuracy,
|
||||
assuming the activation function for the hidden layer is a <b>non-constant, bounded and monotonically-increasing continuous function</b>.
|
||||
Note that the requirements on the activation function only applies to the hidden layer, the output nodes are always
|
||||
assumed to be linear, so as to not restrict the range of output values.
|
||||
|
||||
<p>
|
||||
We note that this theorem is only applicable to a NN with <em>one</em> hidden layer.
|
||||
Therefore, we can easily construct an NN
|
||||
that employs activation functions which do not satisfy the above requirements, as long as we have at least one layer
|
||||
with activation functions that <em>do</em>. Furthermore, although the universal approximation theorem
|
||||
lays the theoretical foundation for regression with neural networks, it does not say anything about how things work in practice:
|
||||
A neural network can still be able to approximate a given function reasonably well without having the flexibility to fit <em>all other</em>
|
||||
functions.
|
||||
Note that the requirements on the activation function only applies to
|
||||
the hidden layer, the output nodes are always assumed to be linear, so
|
||||
as to not restrict the range of output values.
|
||||
|
||||
<p>
|
||||
We note that this theorem is only applicable to an NN with <em>one</em> hidden
|
||||
layer. Therefore, we can easily construct an NN that employs
|
||||
activation functions which do not satisfy the above requirements, as
|
||||
long as we have at least one layer with activation functions that
|
||||
<em>do</em>. Furthermore, although the universal approximation theorem lays
|
||||
the theoretical foundation for regression with neural networks, it
|
||||
does not say anything about how things work in practice: A neural
|
||||
network can still be able to approximate a given function reasonably
|
||||
well without having the flexibility to fit <em>all other</em> functions.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
@@ -319,9 +321,11 @@ $$
|
||||
\end{equation}
|
||||
$$
|
||||
|
||||
In an FFNN of such neurons, the <em>inputs</em> \( x_i \)
|
||||
are the <em>outputs</em> of the neurons in the preceding layer. Furthermore, an MLP is fully-connected,
|
||||
which means that each neuron receives a weighted sum of the outputs of <em>all</em> neurons in the previous layer.
|
||||
<p>
|
||||
In an FFNN of such neurons, the <em>inputs</em> \( x_i \) are the <em>outputs</em> of
|
||||
the neurons in the preceding layer. Furthermore, an MLP is
|
||||
fully-connected, which means that each neuron receives a weighted sum
|
||||
of the outputs of <em>all</em> neurons in the previous layer.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
@@ -348,7 +352,9 @@ $$
|
||||
\end{equation}
|
||||
$$
|
||||
|
||||
where we assume that all nodes in the same layer have identical activation functions, hence the notation \( f_l \)
|
||||
<p>
|
||||
where we assume that all nodes in the same layer have identical
|
||||
activation functions, hence the notation \( f_l \)
|
||||
|
||||
$$
|
||||
\begin{equation}
|
||||
@@ -357,8 +363,11 @@ $$
|
||||
\end{equation}
|
||||
$$
|
||||
|
||||
where \( N_l \) is the number of nodes in layer \( l \). When the output of all the nodes in the first hidden layer are computed,
|
||||
the values of the subsequent layer can be calculated and so forth until the output is obtained.
|
||||
<p>
|
||||
where \( N_l \) is the number of nodes in layer \( l \). When the output of
|
||||
all the nodes in the first hidden layer are computed, the values of
|
||||
the subsequent layer can be calculated and so forth until the output
|
||||
is obtained.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
@@ -395,8 +404,9 @@ $$
|
||||
<h2 id="___sec11">Mathematical model </h2>
|
||||
|
||||
<p>
|
||||
We can generalize this expression to an MLP with \( l \) hidden layers. The complete functional form
|
||||
is,
|
||||
We can generalize this expression to an MLP with \( l \) hidden
|
||||
layers. The complete functional form is,
|
||||
|
||||
$$
|
||||
\begin{align}
|
||||
&y^{l+1}_1\! = \!f_{l+1}\!\left[\!\sum_{j=1}^{N_l}\! w_{1j}^3 f_l\!\left(\!\sum_{k=1}^{N_{l-1}}\! w_{jk}^2 f_{l-1}\!\left(\!
|
||||
@@ -407,7 +417,9 @@ $$
|
||||
\end{align}
|
||||
$$
|
||||
|
||||
which illustrates a basic property of MLPs: The only independent variables are the input values \( x_n \).
|
||||
<p>
|
||||
which illustrates a basic property of MLPs: The only independent
|
||||
variables are the input values \( x_n \).
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
@@ -415,15 +427,18 @@ which illustrates a basic property of MLPs: The only independent variables are t
|
||||
<h2 id="___sec12">Mathematical model </h2>
|
||||
|
||||
<p>
|
||||
This confirms that an MLP,
|
||||
despite its quite convoluted mathematical form, is nothing more than an analytic function, specifically a
|
||||
mapping of real-valued vectors \( \vec{x} \in \mathbb{R}^n \rightarrow \vec{y} \in \mathbb{R}^m \).
|
||||
In our example, \( n=2 \) and \( m=1 \). Consequentially,
|
||||
the number of input and output values of the function we want to fit must be equal to the number of inputs and outputs of our MLP.
|
||||
This confirms that an MLP, despite its quite convoluted mathematical
|
||||
form, is nothing more than an analytic function, specifically a
|
||||
mapping of real-valued vectors \( \vec{x} \in \mathbb{R}^n \rightarrow
|
||||
\vec{y} \in \mathbb{R}^m \). In our example, \( n=2 \) and
|
||||
\( m=1 \). Consequentially, the number of input and output values of the
|
||||
function we want to fit must be equal to the number of inputs and
|
||||
outputs of our MLP.
|
||||
|
||||
<p>
|
||||
Furthermore, the flexibility and universality of a MLP can be illustrated by realizing that
|
||||
the expression is essentially a nested sum of scaled activation functions of the form
|
||||
Furthermore, the flexibility and universality of a MLP can be
|
||||
illustrated by realizing that the expression is essentially a nested
|
||||
sum of scaled activation functions of the form
|
||||
|
||||
$$
|
||||
\begin{equation}
|
||||
@@ -432,15 +447,18 @@ $$
|
||||
\end{equation}
|
||||
$$
|
||||
|
||||
where the parameters \( c_i \) are weights and biases. By adjusting these parameters, the activation functions
|
||||
can be shifted up and down or left and right, change slope or be rescaled
|
||||
which is the key to the flexibility of a neural network.
|
||||
<p>
|
||||
where the parameters \( c_i \) are weights and biases. By adjusting these
|
||||
parameters, the activation functions can be shifted up and down or
|
||||
left and right, change slope or be rescaled which is the key to the
|
||||
flexibility of a neural network.
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h3 id="___sec13">Matrix-vector notation </h3>
|
||||
|
||||
<p>
|
||||
We can introduce a more convenient notation for the activations in a NN.
|
||||
|
||||
<p>
|
||||
@@ -479,6 +497,7 @@ $$
|
||||
|
||||
<h3 id="___sec14">Matrix-vector notation and activation </h3>
|
||||
|
||||
<p>
|
||||
The activation of node \( i \) in layer 2 is
|
||||
|
||||
$$
|
||||
@@ -489,9 +508,11 @@ $$
|
||||
\end{equation}
|
||||
$$
|
||||
|
||||
This is not just a convenient and compact notation, but also
|
||||
a useful and intuitive way to think about MLPs: The output is calculated by a series of matrix-vector multiplications
|
||||
and vector additions that are used as input to the activation functions. For each operation
|
||||
<p>
|
||||
This is not just a convenient and compact notation, but also a useful
|
||||
and intuitive way to think about MLPs: The output is calculated by a
|
||||
series of matrix-vector multiplications and vector additions that are
|
||||
used as input to the activation functions. For each operation
|
||||
\( \mathrm{W}_l \vec{y}_{l-1} \) we move forward one layer.
|
||||
|
||||
<p>
|
||||
@@ -500,9 +521,10 @@ and vector additions that are used as input to the activation functions. For eac
|
||||
<h3 id="___sec15">Activation functions </h3>
|
||||
|
||||
<p>
|
||||
A property that characterizes a neural network, other than its connectivity, is the choice of activation function(s).
|
||||
As described in, the following restrictions are imposed on an activation function for a FFNN
|
||||
to fulfill the universal approximation theorem
|
||||
A property that characterizes a neural network, other than its
|
||||
connectivity, is the choice of activation function(s). As described
|
||||
in, the following restrictions are imposed on an activation function
|
||||
for a FFNN to fulfill the universal approximation theorem
|
||||
|
||||
<ul>
|
||||
<li> Non-constant</li>
|
||||
@@ -516,14 +538,16 @@ to fulfill the universal approximation theorem
|
||||
<h3 id="___sec16">Activation functions, Logistic and Hyperbolic ones </h3>
|
||||
|
||||
<p>
|
||||
The second requirement excludes all linear functions. Furthermore, in a MLP with only linear activation functions, each
|
||||
layer simply performs a linear transformation of its inputs.
|
||||
The second requirement excludes all linear functions. Furthermore, in
|
||||
a MLP with only linear activation functions, each layer simply
|
||||
performs a linear transformation of its inputs.
|
||||
|
||||
<p>
|
||||
Regardless of the number of layers,
|
||||
the output of the NN will be nothing but a linear function of the inputs. Thus we need to introduce some kind of
|
||||
non-linearity to the NN to be able to fit non-linear functions
|
||||
Typical examples are the logistic <em>Sigmoid</em>
|
||||
Regardless of the number of layers, the output of the NN will be
|
||||
nothing but a linear function of the inputs. Thus we need to introduce
|
||||
some kind of non-linearity to the NN to be able to fit non-linear
|
||||
functions Typical examples are the logistic <em>Sigmoid</em>
|
||||
|
||||
$$
|
||||
\begin{equation}
|
||||
f(x) = \frac{1}{1 + e^{-x}},
|
||||
@@ -544,11 +568,12 @@ $$
|
||||
|
||||
<h3 id="___sec17">Relevance </h3>
|
||||
|
||||
The <em>sigmoid</em> function are more biologically plausible because
|
||||
the output of inactive neurons are zero. Such activation function are called <em>one-sided</em>. However,
|
||||
it has been shown that the hyperbolic tangent
|
||||
performs better than the sigmoid for training MLPs.
|
||||
has become the most popular for <em>deep neural networks</em>
|
||||
<p>
|
||||
The <em>sigmoid</em> function are more biologically plausible because the
|
||||
output of inactive neurons are zero. Such activation function are
|
||||
called <em>one-sided</em>. However, it has been shown that the hyperbolic
|
||||
tangent performs better than the sigmoid for training MLPs. has
|
||||
become the most popular for <em>deep neural networks</em>
|
||||
|
||||
<p>
|
||||
|
||||
|
||||
Reference in New Issue
Block a user