cleaning up
This commit is contained in:
@@ -137,16 +137,13 @@ means that each node in the output layer has a linear activation function). The
|
||||
there are no cycles, thus RBFs can be viewed as a type of fully-connected FFNN. They are however usually treated as
|
||||
a separate type of NN due the unusual activation functions.
|
||||
|
||||
|
||||
Other types of NNs could also be mentioned, but are outside the scope of this work. We will now move on to a detailed description
|
||||
of how a fully-connected FFNN works, and how it can be used to interpolate data sets.
|
||||
|
||||
!split
|
||||
===== Multilayer perceptrons =====
|
||||
|
||||
One use often so-called fully-connected feed-forward neural networks with three
|
||||
or more layers (an input layer, one or more hidden layers and an output layer)
|
||||
consisting of neurons that have non-linear activation functions.
|
||||
One uses often so-called fully-connected feed-forward neural networks
|
||||
with three or more layers (an input layer, one or more hidden layers
|
||||
and an output layer) consisting of neurons that have non-linear
|
||||
activation functions.
|
||||
|
||||
Such networks are often called *multilayer perceptrons* (MLPs)
|
||||
|
||||
@@ -156,16 +153,20 @@ Such networks are often called *multilayer perceptrons* (MLPs)
|
||||
According to the *Universal approximation theorem*, a feed-forward neural network with just a single hidden layer containing
|
||||
a finite number of neurons can approximate a continuous multidimensional function to arbitrary accuracy,
|
||||
assuming the activation function for the hidden layer is a _non-constant, bounded and monotonically-increasing continuous function_.
|
||||
Note that the requirements on the activation function only applies to the hidden layer, the output nodes are always
|
||||
assumed to be linear, so as to not restrict the range of output values.
|
||||
|
||||
We note that this theorem is only applicable to a NN with *one* hidden layer.
|
||||
Therefore, we can easily construct an NN
|
||||
that employs activation functions which do not satisfy the above requirements, as long as we have at least one layer
|
||||
with activation functions that *do*. Furthermore, although the universal approximation theorem
|
||||
lays the theoretical foundation for regression with neural networks, it does not say anything about how things work in practice:
|
||||
A neural network can still be able to approximate a given function reasonably well without having the flexibility to fit *all other*
|
||||
functions.
|
||||
Note that the requirements on the activation function only applies to
|
||||
the hidden layer, the output nodes are always assumed to be linear, so
|
||||
as to not restrict the range of output values.
|
||||
|
||||
We note that this theorem is only applicable to an NN with *one* hidden
|
||||
layer. Therefore, we can easily construct an NN that employs
|
||||
activation functions which do not satisfy the above requirements, as
|
||||
long as we have at least one layer with activation functions that
|
||||
*do*. Furthermore, although the universal approximation theorem lays
|
||||
the theoretical foundation for regression with neural networks, it
|
||||
does not say anything about how things work in practice: A neural
|
||||
network can still be able to approximate a given function reasonably
|
||||
well without having the flexibility to fit *all other* functions.
|
||||
|
||||
|
||||
|
||||
@@ -178,9 +179,11 @@ functions.
|
||||
label{artificialNeuron2}
|
||||
\end{equation}
|
||||
!et
|
||||
In an FFNN of such neurons, the *inputs* $x_i$
|
||||
are the *outputs* of the neurons in the preceding layer. Furthermore, an MLP is fully-connected,
|
||||
which means that each neuron receives a weighted sum of the outputs of *all* neurons in the previous layer.
|
||||
|
||||
In an FFNN of such neurons, the *inputs* $x_i$ are the *outputs* of
|
||||
the neurons in the preceding layer. Furthermore, an MLP is
|
||||
fully-connected, which means that each neuron receives a weighted sum
|
||||
of the outputs of *all* neurons in the previous layer.
|
||||
|
||||
!split
|
||||
===== Mathematical model =====
|
||||
@@ -201,7 +204,9 @@ producing the output $y_i^1$ of all neurons in layer 1,
|
||||
label{outputLayer1}
|
||||
\end{equation}
|
||||
!et
|
||||
where we assume that all nodes in the same layer have identical activation functions, hence the notation $f_l$
|
||||
|
||||
where we assume that all nodes in the same layer have identical
|
||||
activation functions, hence the notation $f_l$
|
||||
|
||||
!bt
|
||||
\begin{equation}
|
||||
@@ -209,8 +214,11 @@ where we assume that all nodes in the same layer have identical activation funct
|
||||
label{generalLayer}
|
||||
\end{equation}
|
||||
!et
|
||||
where $N_l$ is the number of nodes in layer $l$. When the output of all the nodes in the first hidden layer are computed,
|
||||
the values of the subsequent layer can be calculated and so forth until the output is obtained.
|
||||
|
||||
where $N_l$ is the number of nodes in layer $l$. When the output of
|
||||
all the nodes in the first hidden layer are computed, the values of
|
||||
the subsequent layer can be calculated and so forth until the output
|
||||
is obtained.
|
||||
|
||||
|
||||
|
||||
@@ -239,8 +247,9 @@ where we have substituted $y_m^1$ with. Finally, the NN output yields,
|
||||
!split
|
||||
===== Mathematical model =====
|
||||
|
||||
We can generalize this expression to an MLP with $l$ hidden layers. The complete functional form
|
||||
is,
|
||||
We can generalize this expression to an MLP with $l$ hidden
|
||||
layers. The complete functional form is,
|
||||
|
||||
!bt
|
||||
\begin{align}
|
||||
&y^{l+1}_1\! = \!f_{l+1}\!\left[\!\sum_{j=1}^{N_l}\! w_{1j}^3 f_l\!\left(\!\sum_{k=1}^{N_{l-1}}\! w_{jk}^2 f_{l-1}\!\left(\!
|
||||
@@ -250,31 +259,39 @@ is,
|
||||
label{completeNN}
|
||||
\end{align}
|
||||
!et
|
||||
which illustrates a basic property of MLPs: The only independent variables are the input values $x_n$.
|
||||
|
||||
which illustrates a basic property of MLPs: The only independent
|
||||
variables are the input values $x_n$.
|
||||
|
||||
!split
|
||||
===== Mathematical model =====
|
||||
|
||||
This confirms that an MLP,
|
||||
despite its quite convoluted mathematical form, is nothing more than an analytic function, specifically a
|
||||
mapping of real-valued vectors $\vec{x} \in \mathbb{R}^n \rightarrow \vec{y} \in \mathbb{R}^m$.
|
||||
In our example, $n=2$ and $m=1$. Consequentially,
|
||||
the number of input and output values of the function we want to fit must be equal to the number of inputs and outputs of our MLP.
|
||||
This confirms that an MLP, despite its quite convoluted mathematical
|
||||
form, is nothing more than an analytic function, specifically a
|
||||
mapping of real-valued vectors $\vec{x} \in \mathbb{R}^n \rightarrow
|
||||
\vec{y} \in \mathbb{R}^m$. In our example, $n=2$ and
|
||||
$m=1$. Consequentially, the number of input and output values of the
|
||||
function we want to fit must be equal to the number of inputs and
|
||||
outputs of our MLP.
|
||||
|
||||
Furthermore, the flexibility and universality of a MLP can be illustrated by realizing that
|
||||
the expression is essentially a nested sum of scaled activation functions of the form
|
||||
Furthermore, the flexibility and universality of a MLP can be
|
||||
illustrated by realizing that the expression is essentially a nested
|
||||
sum of scaled activation functions of the form
|
||||
|
||||
!bt
|
||||
\begin{equation}
|
||||
h(x) = c_1 f(c_2 x + c_3) + c_4
|
||||
\end{equation}
|
||||
!et
|
||||
where the parameters $c_i$ are weights and biases. By adjusting these parameters, the activation functions
|
||||
can be shifted up and down or left and right, change slope or be rescaled
|
||||
which is the key to the flexibility of a neural network.
|
||||
|
||||
where the parameters $c_i$ are weights and biases. By adjusting these
|
||||
parameters, the activation functions can be shifted up and down or
|
||||
left and right, change slope or be rescaled which is the key to the
|
||||
flexibility of a neural network.
|
||||
|
||||
!split
|
||||
=== Matrix-vector notation ===
|
||||
|
||||
We can introduce a more convenient notation for the activations in a NN.
|
||||
|
||||
Additionally, we can represent the biases and activations
|
||||
@@ -307,6 +324,7 @@ the equation for the activations of hidden layer 2 in
|
||||
|
||||
!split
|
||||
=== Matrix-vector notation and activation ===
|
||||
|
||||
The activation of node $i$ in layer 2 is
|
||||
|
||||
!bt
|
||||
@@ -315,19 +333,22 @@ The activation of node $i$ in layer 2 is
|
||||
f_2\left(\sum_{j=1}^3 w^2_{ij} y_j^1 + b^2_i\right).
|
||||
\end{equation}
|
||||
!et
|
||||
This is not just a convenient and compact notation, but also
|
||||
a useful and intuitive way to think about MLPs: The output is calculated by a series of matrix-vector multiplications
|
||||
and vector additions that are used as input to the activation functions. For each operation
|
||||
$\mathrm{W}_l \vec{y}_{l-1}$ we move forward one layer.
|
||||
|
||||
This is not just a convenient and compact notation, but also a useful
|
||||
and intuitive way to think about MLPs: The output is calculated by a
|
||||
series of matrix-vector multiplications and vector additions that are
|
||||
used as input to the activation functions. For each operation
|
||||
$\mathrm{W}_l \vec{y}_{l-1}$ we move forward one layer.
|
||||
|
||||
|
||||
!split
|
||||
=== Activation functions ===
|
||||
|
||||
|
||||
A property that characterizes a neural network, other than its connectivity, is the choice of activation function(s).
|
||||
As described in, the following restrictions are imposed on an activation function for a FFNN
|
||||
to fulfill the universal approximation theorem
|
||||
A property that characterizes a neural network, other than its
|
||||
connectivity, is the choice of activation function(s). As described
|
||||
in, the following restrictions are imposed on an activation function
|
||||
for a FFNN to fulfill the universal approximation theorem
|
||||
|
||||
* Non-constant
|
||||
|
||||
@@ -340,13 +361,15 @@ to fulfill the universal approximation theorem
|
||||
!split
|
||||
=== Activation functions, Logistic and Hyperbolic ones ===
|
||||
|
||||
The second requirement excludes all linear functions. Furthermore, in a MLP with only linear activation functions, each
|
||||
layer simply performs a linear transformation of its inputs.
|
||||
The second requirement excludes all linear functions. Furthermore, in
|
||||
a MLP with only linear activation functions, each layer simply
|
||||
performs a linear transformation of its inputs.
|
||||
|
||||
Regardless of the number of layers, the output of the NN will be
|
||||
nothing but a linear function of the inputs. Thus we need to introduce
|
||||
some kind of non-linearity to the NN to be able to fit non-linear
|
||||
functions Typical examples are the logistic *Sigmoid*
|
||||
|
||||
Regardless of the number of layers,
|
||||
the output of the NN will be nothing but a linear function of the inputs. Thus we need to introduce some kind of
|
||||
non-linearity to the NN to be able to fit non-linear functions
|
||||
Typical examples are the logistic *Sigmoid*
|
||||
!bt
|
||||
\begin{equation}
|
||||
f(x) = \frac{1}{1 + e^{-x}},
|
||||
@@ -363,11 +386,12 @@ and the *hyperbolic tangent* function
|
||||
|
||||
!split
|
||||
=== Relevance ===
|
||||
The *sigmoid* function are more biologically plausible because
|
||||
the output of inactive neurons are zero. Such activation function are called *one-sided*. However,
|
||||
it has been shown that the hyperbolic tangent
|
||||
performs better than the sigmoid for training MLPs.
|
||||
has become the most popular for *deep neural networks*
|
||||
|
||||
The *sigmoid* function are more biologically plausible because the
|
||||
output of inactive neurons are zero. Such activation function are
|
||||
called *one-sided*. However, it has been shown that the hyperbolic
|
||||
tangent performs better than the sigmoid for training MLPs. has
|
||||
become the most popular for *deep neural networks*
|
||||
|
||||
!bc pycod
|
||||
"""The sigmoid function (or the logistic curve) is a
|
||||
|
||||
Reference in New Issue
Block a user