1087 lines
44 KiB
Plaintext
1087 lines
44 KiB
Plaintext
{
|
|
"cells": [
|
|
{
|
|
"cell_type": "markdown",
|
|
"metadata": {},
|
|
"source": [
|
|
"<!-- dom:TITLE: Data Analysis and Machine Learning: Elements of machine learning -->\n",
|
|
"# Data Analysis and Machine Learning: Elements of machine learning\n",
|
|
"<!-- dom:AUTHOR: Morten Hjorth-Jensen at Department of Physics, University of Oslo & Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University -->\n",
|
|
"<!-- Author: --> \n",
|
|
"**Morten Hjorth-Jensen**, Department of Physics, University of Oslo and Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University\n",
|
|
"\n",
|
|
"Date: **May 30, 2018**\n",
|
|
"\n",
|
|
"Copyright 1999-2018, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license\n",
|
|
"\n",
|
|
"\n",
|
|
"\n",
|
|
"## What is Machine Learning?\n",
|
|
"\n",
|
|
"Machine learning is the science of giving computers the ability to\n",
|
|
"learn without being explicitly programmed. The idea is that there\n",
|
|
"exist generic algorithms which can be used to find patterns in a broad\n",
|
|
"class of data sets without having to write code specifically for each\n",
|
|
"problem. The algorithm will build its own logic based on the data.\n",
|
|
"\n",
|
|
"Machine learning is a subfield of computer science, and is closely\n",
|
|
"related to computational statistics. It evolved from the study of\n",
|
|
"pattern recognition in artificial intelligence (AI) research, and has\n",
|
|
"made contributions to AI tasks like computer vision, natural language\n",
|
|
"processing and speech recognition. It has also, especially in later\n",
|
|
"years, found applications in a wide variety of other areas, including\n",
|
|
"bioinformatics, economy, physics, finance and marketing.\n",
|
|
"\n",
|
|
"## Types of Machine Learning\n",
|
|
"\n",
|
|
"\n",
|
|
"The approaches to machine learning are many, but are often split into two main categories. \n",
|
|
"In *supervised learning* we know the answer to a problem,\n",
|
|
"and let the computer deduce the logic behind it. On the other hand, *unsupervised learning*\n",
|
|
"is a method for finding patterns and relationship in data sets without any prior knowledge of the system.\n",
|
|
"Some authours also operate with a third category, namely *reinforcement learning*. This is a paradigm \n",
|
|
"of learning inspired by behavioural psychology, where learning is achieved by trial-and-error, \n",
|
|
"solely from rewards and punishment.\n",
|
|
"\n",
|
|
"Another way to categorize machine learning tasks is to consider the desired output of a system.\n",
|
|
"Some of the most common tasks are:\n",
|
|
"\n",
|
|
" * Classification: Outputs are divided into two or more classes. The goal is to produce a model that assigns inputs into one of these classes. An example is to identify digits based on pictures of hand-written ones. Classification is typically supervised learning.\n",
|
|
"\n",
|
|
" * Regression: Finding a functional relationship between an input data set and a reference data set. The goal is to construct a function that maps input data to continuous output values.\n",
|
|
"\n",
|
|
" * Clustering: Data are divided into groups with certain common traits, without knowing the different groups beforehand. It is thus a form of unsupervised learning.\n",
|
|
"\n",
|
|
"## Artificial neurons\n",
|
|
"\n",
|
|
"The field of artificial neural networks has a long history of\n",
|
|
"development, and is closely connected with the advancement of computer\n",
|
|
"science and computers in general. A model of artificial neurons was\n",
|
|
"first developed by McCulloch and Pitts in 1943 to study signal\n",
|
|
"processing in the brain and has later been refined by others. The\n",
|
|
"general idea is to mimic neural networks in the human brain, which is\n",
|
|
"composed of billions of neurons that communicate with each other by\n",
|
|
"sending electrical signals. Each neuron accumulates its incoming\n",
|
|
"signals, which must exceed an activation threshold to yield an\n",
|
|
"output. If the threshold is not overcome, the neuron remains inactive,\n",
|
|
"i.e. has zero output.\n",
|
|
"\n",
|
|
"This behaviour has inspired a simple mathematical model for an artificial neuron."
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"metadata": {},
|
|
"source": [
|
|
"<!-- Equation labels as ordinary links -->\n",
|
|
"<div id=\"artificialNeuron\"></div>\n",
|
|
"\n",
|
|
"$$\n",
|
|
"\\begin{equation}\n",
|
|
" y = f\\left(\\sum_{i=1}^n w_ix_i\\right) = f(u)\n",
|
|
"\\label{artificialNeuron} \\tag{1}\n",
|
|
"\\end{equation}\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"metadata": {},
|
|
"source": [
|
|
"Here, the output $y$ of the neuron is the value of its activation function, which have as input\n",
|
|
"a weighted sum of signals $x_i, \\dots ,x_n$ received by $n$ other neurons.\n",
|
|
"\n",
|
|
"Conceptually, it is helpful to divide neural networks into four\n",
|
|
"categories:\n",
|
|
"1. general purpose neural networks for supervised learning,\n",
|
|
"\n",
|
|
"2. neural networks designed specifically for image processing, the most prominent example of this class being Convolutional Neural Networks (CNNs),\n",
|
|
"\n",
|
|
"3. neural networks for sequential data such as Recurrent Neural Networks (RNNs), and\n",
|
|
"\n",
|
|
"4. neural networks for unsupervised learning such as Deep Boltzmann Machines.\n",
|
|
"\n",
|
|
"In physics, DNNs and CNNs have already found numerous applications. In\n",
|
|
"statistical physics, they have been applied to detect phase\n",
|
|
"transitions in 2D Ising and Potts models, lattice gauge theories, and\n",
|
|
"different phases of polymers.\n",
|
|
"Deep learning has also found interesting applications in quantum\n",
|
|
"physics. Various quantum phase transitions can be detected and studied\n",
|
|
"using DNNs and CNNs, including the transverse-field Ising model,\n",
|
|
"topological phases, and even non-equilibrium many-body\n",
|
|
"localization. Representing quantum states as DNNs quantum state\n",
|
|
"tomography are among some of the impressive\n",
|
|
"achievements to reveal the potential of DNNs to facilitate the study\n",
|
|
"of quantum systems.\n",
|
|
"\n",
|
|
"In quantum information theory, it has been shown that one can perform\n",
|
|
"gate decompositions with the help of neural. In lattice quantum chromodynamics,\n",
|
|
"DNNs have been used to learn action parameters in regions of parameter\n",
|
|
"space where PCA fails. Last but not least,\n",
|
|
"DNNs also found place in the study of quantum, and in scattering theory to learn\n",
|
|
"$s$-wave scattering length of potentials.\n",
|
|
"\n",
|
|
"## Neural network types\n",
|
|
"\n",
|
|
"An artificial neural network (NN), is a computational model that\n",
|
|
"consists of layers of connected neurons, or *nodes*. It is supposed\n",
|
|
"to mimic a biological nervous system by letting each neuron interact\n",
|
|
"with other neurons by sending signals in the form of mathematical\n",
|
|
"functions between layers. A wide variety of different NNs have been\n",
|
|
"developed, but most of them consist of an input layer, an output layer\n",
|
|
"and eventual layers in-between, called *hidden layers*. All layers can\n",
|
|
"contain an arbitrary number of nodes, and each connection between two\n",
|
|
"nodes is associated with a weight variable.\n",
|
|
"\n",
|
|
"Neural networks (also called neural nets) are neural-inspired\n",
|
|
"nonlinear models for supervised learning. As we will see, neural nets\n",
|
|
"can be viewed as natural, more powerful extensions of supervised\n",
|
|
"learning methods such as linear and logistic regression and soft-max\n",
|
|
"methods.\n",
|
|
"\n",
|
|
"\n",
|
|
"## Feed-forward neural networks\n",
|
|
"\n",
|
|
"The feed-forward neural network (FFNN) was the first and simplest type of NN devised. In this network, \n",
|
|
"the information moves in only one direction: forward through the layers.\n",
|
|
"\n",
|
|
"Nodes are represented by circles, while the arrows display the connections between the nodes, including the \n",
|
|
"direction of information flow. Additionally, each arrow corresponds to a weight variable, not displayed here. \n",
|
|
"We observe that each node in a layer is connected to *all* nodes in the subsequent layer, \n",
|
|
"making this a so-called *fully-connected* FFNN. \n",
|
|
"\n",
|
|
"\n",
|
|
"\n",
|
|
"A different variant of FFNNs are *convolutional neural networks* (CNNs), which have a connectivity pattern\n",
|
|
"inspired by the animal visual cortex. Individual neurons in the visual cortex only respond to stimuli from\n",
|
|
"small sub-regions of the visual field, called a receptive field. This makes the neurons well-suited to exploit the strong\n",
|
|
"spatially local correlation present in natural images. The response of each neuron can be approximated mathematically \n",
|
|
"as a convolution operation. \n",
|
|
"\n",
|
|
"CNNs emulate the behaviour of neurons in the visual cortex by enforcing a *local* connectivity pattern\n",
|
|
"between nodes of adjacent layers: Each node\n",
|
|
"in a convolutional layer is connected only to a subset of the nodes in the previous layer, \n",
|
|
"in contrast to the fully-connected FFNN.\n",
|
|
"Often, CNNs \n",
|
|
"consist of several convolutional layers that learn local features of the input, with a fully-connected layer at the end, \n",
|
|
"which gathers all the local data and produces the outputs. They have wide applications in image and video recognition\n",
|
|
"\n",
|
|
"## Recurrent neural networks\n",
|
|
"\n",
|
|
"So far we have only mentioned NNs where information flows in one direction: forward. *Recurrent neural networks* on\n",
|
|
"the other hand, have connections between nodes that form directed *cycles*. This creates a form of \n",
|
|
"internal memory which are able to capture information on what has been calculated before; the output is dependent \n",
|
|
"on the previous computations. Recurrent NNs make use of sequential information by performing the same task for \n",
|
|
"every element in a sequence, where each element depends on previous elements. An example of such information is \n",
|
|
"sentences, making recurrent NNs especially well-suited for handwriting and speech recognition.\n",
|
|
"\n",
|
|
"## Other types of networks\n",
|
|
"\n",
|
|
"There are many other kinds of NNs that have been developed. One type that is specifically designed for interpolation\n",
|
|
"in multidimensional space is the radial basis function (RBF) network. RBFs are typically made up of three layers: \n",
|
|
"an input layer, a hidden layer with non-linear radial symmetric activation functions and a linear output layer (''linear'' here\n",
|
|
"means that each node in the output layer has a linear activation function). The layers are normally fully-connected and \n",
|
|
"there are no cycles, thus RBFs can be viewed as a type of fully-connected FFNN. They are however usually treated as\n",
|
|
"a separate type of NN due the unusual activation functions.\n",
|
|
"\n",
|
|
"\n",
|
|
"Other types of NNs could also be mentioned, but are outside the scope of this work. We will now move on to a detailed description\n",
|
|
"of how a fully-connected FFNN works, and how it can be used to interpolate data sets. \n",
|
|
"\n",
|
|
"## Multilayer perceptrons\n",
|
|
"\n",
|
|
"One use often so-called fully-connected feed-forward neural networks with three\n",
|
|
"or more layers (an input layer, one or more hidden layers and an output layer)\n",
|
|
"consisting of neurons that have non-linear activation functions.\n",
|
|
"\n",
|
|
"Such networks are often called *multilayer perceptrons* (MLPs)\n",
|
|
"\n",
|
|
"## Why multilayer perceptrons?\n",
|
|
"\n",
|
|
"According to the *Universal approximation theorem*, a feed-forward neural network with just a single hidden layer containing \n",
|
|
"a finite number of neurons can approximate a continuous multidimensional function to arbitrary accuracy, \n",
|
|
"assuming the activation function for the hidden layer is a **non-constant, bounded and monotonically-increasing continuous function**.\n",
|
|
"Note that the requirements on the activation function only applies to the hidden layer, the output nodes are always\n",
|
|
"assumed to be linear, so as to not restrict the range of output values. \n",
|
|
"\n",
|
|
"We note that this theorem is only applicable to a NN with *one* hidden layer. \n",
|
|
"Therefore, we can easily construct an NN \n",
|
|
"that employs activation functions which do not satisfy the above requirements, as long as we have at least one layer\n",
|
|
"with activation functions that *do*. Furthermore, although the universal approximation theorem\n",
|
|
"lays the theoretical foundation for regression with neural networks, it does not say anything about how things work in practice: \n",
|
|
"A neural network can still be able to approximate a given function reasonably well without having the flexibility to fit *all other*\n",
|
|
"functions. \n",
|
|
"\n",
|
|
"\n",
|
|
"\n",
|
|
"## Mathematical model"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"metadata": {},
|
|
"source": [
|
|
"<!-- Equation labels as ordinary links -->\n",
|
|
"<div id=\"artificialNeuron2\"></div>\n",
|
|
"\n",
|
|
"$$\n",
|
|
"\\begin{equation}\n",
|
|
" y = f\\left(\\sum_{i=1}^n w_ix_i + b_i\\right) = f(u)\n",
|
|
"\\label{artificialNeuron2} \\tag{2}\n",
|
|
"\\end{equation}\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"metadata": {},
|
|
"source": [
|
|
"In an FFNN of such neurons, the *inputs* $x_i$\n",
|
|
"are the *outputs* of the neurons in the preceding layer. Furthermore, an MLP is fully-connected, \n",
|
|
"which means that each neuron receives a weighted sum of the outputs of *all* neurons in the previous layer. \n",
|
|
"\n",
|
|
"## Mathematical model\n",
|
|
"\n",
|
|
"First, for each node $i$ in the first hidden layer, we calculate a weighted sum $u_i^1$ of the input coordinates $x_j$,"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"metadata": {},
|
|
"source": [
|
|
"<!-- Equation labels as ordinary links -->\n",
|
|
"<div id=\"_auto1\"></div>\n",
|
|
"\n",
|
|
"$$\n",
|
|
"\\begin{equation}\n",
|
|
" u_i^1 = \\sum_{j=1}^2 w_{ij}^1 x_j + b_i^1 \n",
|
|
"\\label{_auto1} \\tag{3}\n",
|
|
"\\end{equation}\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"metadata": {},
|
|
"source": [
|
|
"This value is the argument to the activation function $f_1$ of each neuron $i$,\n",
|
|
"producing the output $y_i^1$ of all neurons in layer 1,"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"metadata": {},
|
|
"source": [
|
|
"<!-- Equation labels as ordinary links -->\n",
|
|
"<div id=\"outputLayer1\"></div>\n",
|
|
"\n",
|
|
"$$\n",
|
|
"\\begin{equation}\n",
|
|
" y_i^1 = f_1(u_i^1) = f_1\\left(\\sum_{j=1}^2 w_{ij}^1 x_j + b_i^1\\right)\n",
|
|
"\\label{outputLayer1} \\tag{4}\n",
|
|
"\\end{equation}\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"metadata": {},
|
|
"source": [
|
|
"where we assume that all nodes in the same layer have identical activation functions, hence the notation $f_l$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"metadata": {},
|
|
"source": [
|
|
"<!-- Equation labels as ordinary links -->\n",
|
|
"<div id=\"generalLayer\"></div>\n",
|
|
"\n",
|
|
"$$\n",
|
|
"\\begin{equation}\n",
|
|
" y_i^l = f_l(u_i^l) = f_l\\left(\\sum_{j=1}^{N_{l-1}} w_{ij}^l y_j^{l-1} + b_i^l\\right)\n",
|
|
"\\label{generalLayer} \\tag{5}\n",
|
|
"\\end{equation}\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"metadata": {},
|
|
"source": [
|
|
"where $N_l$ is the number of nodes in layer $l$. When the output of all the nodes in the first hidden layer are computed,\n",
|
|
"the values of the subsequent layer can be calculated and so forth until the output is obtained. \n",
|
|
"\n",
|
|
"\n",
|
|
"\n",
|
|
"## Mathematical model\n",
|
|
"\n",
|
|
"The output of neuron $i$ in layer 2 is thus,"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"metadata": {},
|
|
"source": [
|
|
"<!-- Equation labels as ordinary links -->\n",
|
|
"<div id=\"_auto2\"></div>\n",
|
|
"\n",
|
|
"$$\n",
|
|
"\\begin{equation}\n",
|
|
" y_i^2 = f_2\\left(\\sum_{j=1}^3 w_{ij}^2 y_j^1 + b_i^2\\right) \n",
|
|
"\\label{_auto2} \\tag{6}\n",
|
|
"\\end{equation}\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"metadata": {},
|
|
"source": [
|
|
"<!-- Equation labels as ordinary links -->\n",
|
|
"<div id=\"outputLayer2\"></div>\n",
|
|
"\n",
|
|
"$$\n",
|
|
"\\begin{equation} \n",
|
|
" = f_2\\left[\\sum_{j=1}^3 w_{ij}^2f_1\\left(\\sum_{k=1}^2 w_{jk}^1 x_k + b_j^1\\right) + b_i^2\\right]\n",
|
|
"\\label{outputLayer2} \\tag{7}\n",
|
|
"\\end{equation}\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"metadata": {},
|
|
"source": [
|
|
"where we have substituted $y_m^1$ with. Finally, the NN output yields,"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"metadata": {},
|
|
"source": [
|
|
"<!-- Equation labels as ordinary links -->\n",
|
|
"<div id=\"_auto3\"></div>\n",
|
|
"\n",
|
|
"$$\n",
|
|
"\\begin{equation}\n",
|
|
" y_1^3 = f_3\\left(\\sum_{j=1}^3 w_{1m}^3 y_j^2 + b_1^3\\right) \n",
|
|
"\\label{_auto3} \\tag{8}\n",
|
|
"\\end{equation}\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"metadata": {},
|
|
"source": [
|
|
"<!-- Equation labels as ordinary links -->\n",
|
|
"<div id=\"_auto4\"></div>\n",
|
|
"\n",
|
|
"$$\n",
|
|
"\\begin{equation} \n",
|
|
" = f_3\\left[\\sum_{j=1}^3 w_{1j}^3 f_2\\left(\\sum_{k=1}^3 w_{jk}^2 f_1\\left(\\sum_{m=1}^2 w_{km}^1 x_m + b_k^1\\right) + b_j^2\\right)\n",
|
|
" + b_1^3\\right]\n",
|
|
"\\label{_auto4} \\tag{9}\n",
|
|
"\\end{equation}\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"metadata": {},
|
|
"source": [
|
|
"## Mathematical model\n",
|
|
"\n",
|
|
"We can generalize this expression to an MLP with $l$ hidden layers. The complete functional form\n",
|
|
"is,"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"metadata": {},
|
|
"source": [
|
|
"<!-- Equation labels as ordinary links -->\n",
|
|
"<div id=\"completeNN\"></div>\n",
|
|
"\n",
|
|
"$$\n",
|
|
"\\begin{equation}\n",
|
|
"y^{l+1}_1\\! = \\!f_{l+1}\\!\\left[\\!\\sum_{j=1}^{N_l}\\! w_{1j}^3 f_l\\!\\left(\\!\\sum_{k=1}^{N_{l-1}}\\! w_{jk}^2 f_{l-1}\\!\\left(\\!\n",
|
|
" \\dots \\!f_1\\!\\left(\\!\\sum_{n=1}^{N_0} \\!w_{mn}^1 x_n\\! + \\!b_m^1\\!\\right)\n",
|
|
" \\!\\dots \\!\\right) \\!+ \\!b_k^2\\!\\right)\n",
|
|
" \\!+ \\!b_1^3\\!\\right] \n",
|
|
"\\label{completeNN} \\tag{10}\n",
|
|
"\\end{equation}\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"metadata": {},
|
|
"source": [
|
|
"which illustrates a basic property of MLPs: The only independent variables are the input values $x_n$. \n",
|
|
"\n",
|
|
"## Mathematical model\n",
|
|
"\n",
|
|
"This confirms that an MLP,\n",
|
|
"despite its quite convoluted mathematical form, is nothing more than an analytic function, specifically a \n",
|
|
"mapping of real-valued vectors $\\vec{x} \\in \\mathbb{R}^n \\rightarrow \\vec{y} \\in \\mathbb{R}^m$. \n",
|
|
"In our example, $n=2$ and $m=1$. Consequentially, \n",
|
|
"the number of input and output values of the function we want to fit must be equal to the number of inputs and outputs of our MLP. \n",
|
|
"\n",
|
|
"Furthermore, the flexibility and universality of a MLP can be illustrated by realizing that \n",
|
|
"the expression is essentially a nested sum of scaled activation functions of the form"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"metadata": {},
|
|
"source": [
|
|
"<!-- Equation labels as ordinary links -->\n",
|
|
"<div id=\"_auto5\"></div>\n",
|
|
"\n",
|
|
"$$\n",
|
|
"\\begin{equation}\n",
|
|
" h(x) = c_1 f(c_2 x + c_3) + c_4\n",
|
|
"\\label{_auto5} \\tag{11}\n",
|
|
"\\end{equation}\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"metadata": {},
|
|
"source": [
|
|
"where the parameters $c_i$ are weights and biases. By adjusting these parameters, the activation functions\n",
|
|
"can be shifted up and down or left and right, change slope or be rescaled \n",
|
|
"which is the key to the flexibility of a neural network. \n",
|
|
"\n",
|
|
"### Matrix-vector notation\n",
|
|
"\n",
|
|
"We can introduce a more convenient notation for the activations in a NN. \n",
|
|
"\n",
|
|
"Additionally, we can represent the biases and activations\n",
|
|
"as layer-wise column vectors $\\vec{b}_l$ and $\\vec{y}_l$, so that the $i$-th element of each vector \n",
|
|
"is the bias $b_i^l$ and activation $y_i^l$ of node $i$ in layer $l$ respectively. \n",
|
|
"\n",
|
|
"We have that $\\mathrm{W}_l$ is a $N_{l-1} \\times N_l$ matrix, while $\\vec{b}_l$ and $\\vec{y}_l$ are $N_l \\times 1$ column vectors. \n",
|
|
"With this notation, the sum in becomes a matrix-vector multiplication, and we can write\n",
|
|
"the equation for the activations of hidden layer 2 in"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"metadata": {},
|
|
"source": [
|
|
"<!-- Equation labels as ordinary links -->\n",
|
|
"<div id=\"_auto6\"></div>\n",
|
|
"\n",
|
|
"$$\n",
|
|
"\\begin{equation}\n",
|
|
" \\vec{y}_2 = f_2(\\mathrm{W}_2 \\vec{y}_{1} + \\vec{b}_{2}) = \n",
|
|
" f_2\\left(\\left[\\begin{array}{ccc}\n",
|
|
" w^2_{11} &w^2_{12} &w^2_{13} \\\\\n",
|
|
" w^2_{21} &w^2_{22} &w^2_{23} \\\\\n",
|
|
" w^2_{31} &w^2_{32} &w^2_{33} \\\\\n",
|
|
" \\end{array} \\right] \\cdot\n",
|
|
" \\left[\\begin{array}{c}\n",
|
|
" y^1_1 \\\\\n",
|
|
" y^1_2 \\\\\n",
|
|
" y^1_3 \\\\\n",
|
|
" \\end{array}\\right] + \n",
|
|
" \\left[\\begin{array}{c}\n",
|
|
" b^2_1 \\\\\n",
|
|
" b^2_2 \\\\\n",
|
|
" b^2_3 \\\\\n",
|
|
" \\end{array}\\right]\\right).\n",
|
|
"\\label{_auto6} \\tag{12}\n",
|
|
"\\end{equation}\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"metadata": {},
|
|
"source": [
|
|
"### Matrix-vector notation and activation\n",
|
|
"\n",
|
|
"The activation of node $i$ in layer 2 is"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"metadata": {},
|
|
"source": [
|
|
"<!-- Equation labels as ordinary links -->\n",
|
|
"<div id=\"_auto7\"></div>\n",
|
|
"\n",
|
|
"$$\n",
|
|
"\\begin{equation}\n",
|
|
" y^2_i = f_2\\Bigr(w^2_{i1}y^1_1 + w^2_{i2}y^1_2 + w^2_{i3}y^1_3 + b^2_i\\Bigr) = \n",
|
|
" f_2\\left(\\sum_{j=1}^3 w^2_{ij} y_j^1 + b^2_i\\right).\n",
|
|
"\\label{_auto7} \\tag{13}\n",
|
|
"\\end{equation}\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"metadata": {},
|
|
"source": [
|
|
"This is not just a convenient and compact notation, but also \n",
|
|
"a useful and intuitive way to think about MLPs: The output is calculated by a series of matrix-vector multiplications\n",
|
|
"and vector additions that are used as input to the activation functions. For each operation \n",
|
|
"$\\mathrm{W}_l \\vec{y}_{l-1}$ we move forward one layer. \n",
|
|
"\n",
|
|
"\n",
|
|
"### Activation functions\n",
|
|
"\n",
|
|
"A property that characterizes a neural network, other than its connectivity, is the choice of activation function(s). \n",
|
|
"As described in, the following restrictions are imposed on an activation function for a FFNN\n",
|
|
"to fulfill the universal approximation theorem\n",
|
|
"\n",
|
|
" * Non-constant\n",
|
|
"\n",
|
|
" * Bounded\n",
|
|
"\n",
|
|
" * Monotonically-increasing\n",
|
|
"\n",
|
|
" * Continuous\n",
|
|
"\n",
|
|
"### Activation functions, Logistic and Hyperbolic ones\n",
|
|
"\n",
|
|
"The second requirement excludes all linear functions. Furthermore, in a MLP with only linear activation functions, each \n",
|
|
"layer simply performs a linear transformation of its inputs.\n",
|
|
"\n",
|
|
"Regardless of the number of layers, \n",
|
|
"the output of the NN will be nothing but a linear function of the inputs. Thus we need to introduce some kind of \n",
|
|
"non-linearity to the NN to be able to fit non-linear functions\n",
|
|
"Typical examples are the logistic *Sigmoid*"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"metadata": {},
|
|
"source": [
|
|
"<!-- Equation labels as ordinary links -->\n",
|
|
"<div id=\"sigmoidActivationFunction\"></div>\n",
|
|
"\n",
|
|
"$$\n",
|
|
"\\begin{equation}\n",
|
|
" f(x) = \\frac{1}{1 + e^{-x}},\n",
|
|
"\\label{sigmoidActivationFunction} \\tag{14}\n",
|
|
"\\end{equation}\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"metadata": {},
|
|
"source": [
|
|
"and the *hyperbolic tangent* function"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"metadata": {},
|
|
"source": [
|
|
"<!-- Equation labels as ordinary links -->\n",
|
|
"<div id=\"tanhActivationFunction\"></div>\n",
|
|
"\n",
|
|
"$$\n",
|
|
"\\begin{equation}\n",
|
|
" f(x) = \\tanh(x)\n",
|
|
"\\label{tanhActivationFunction} \\tag{15}\n",
|
|
"\\end{equation}\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"metadata": {},
|
|
"source": [
|
|
"### Relevance\n",
|
|
"\n",
|
|
"The *sigmoid* function are more biologically plausible because \n",
|
|
"the output of inactive neurons are zero. Such activation function are called *one-sided*. However,\n",
|
|
"it has been shown that the hyperbolic tangent \n",
|
|
"performs better than the sigmoid for training MLPs. \n",
|
|
"has become the most popular for *deep neural networks*"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "code",
|
|
"execution_count": 1,
|
|
"metadata": {
|
|
"collapsed": false
|
|
},
|
|
"outputs": [],
|
|
"source": [
|
|
"%matplotlib inline\n",
|
|
"\n",
|
|
"\"\"\"The sigmoid function (or the logistic curve) is a \n",
|
|
"function that takes any real number, z, and outputs a number (0,1).\n",
|
|
"It is useful in neural networks for assigning weights on a relative scale.\n",
|
|
"The value z is the weighted sum of parameters involved in the learning algorithm.\"\"\"\n",
|
|
"\n",
|
|
"import numpy\n",
|
|
"import matplotlib.pyplot as plt\n",
|
|
"import math as mt\n",
|
|
"\n",
|
|
"z = numpy.arange(-5, 5, .1)\n",
|
|
"sigma_fn = numpy.vectorize(lambda z: 1/(1+numpy.exp(-z)))\n",
|
|
"sigma = sigma_fn(z)\n",
|
|
"\n",
|
|
"fig = plt.figure()\n",
|
|
"ax = fig.add_subplot(111)\n",
|
|
"ax.plot(z, sigma)\n",
|
|
"ax.set_ylim([-0.1, 1.1])\n",
|
|
"ax.set_xlim([-5,5])\n",
|
|
"ax.grid(True)\n",
|
|
"ax.set_xlabel('z')\n",
|
|
"ax.set_title('sigmoid function')\n",
|
|
"\n",
|
|
"plt.show()\n",
|
|
"\n",
|
|
"\"\"\"Step Function\"\"\"\n",
|
|
"z = numpy.arange(-5, 5, .02)\n",
|
|
"step_fn = numpy.vectorize(lambda z: 1.0 if z >= 0.0 else 0.0)\n",
|
|
"step = step_fn(z)\n",
|
|
"\n",
|
|
"fig = plt.figure()\n",
|
|
"ax = fig.add_subplot(111)\n",
|
|
"ax.plot(z, step)\n",
|
|
"ax.set_ylim([-0.5, 1.5])\n",
|
|
"ax.set_xlim([-5,5])\n",
|
|
"ax.grid(True)\n",
|
|
"ax.set_xlabel('z')\n",
|
|
"ax.set_title('step function')\n",
|
|
"\n",
|
|
"plt.show()\n",
|
|
"\n",
|
|
"\"\"\"Sine Function\"\"\"\n",
|
|
"z = numpy.arange(-2*mt.pi, 2*mt.pi, 0.1)\n",
|
|
"t = numpy.sin(z)\n",
|
|
"\n",
|
|
"fig = plt.figure()\n",
|
|
"ax = fig.add_subplot(111)\n",
|
|
"ax.plot(z, t)\n",
|
|
"ax.set_ylim([-1.0, 1.0])\n",
|
|
"ax.set_xlim([-2*mt.pi,2*mt.pi])\n",
|
|
"ax.grid(True)\n",
|
|
"ax.set_xlabel('z')\n",
|
|
"ax.set_title('sine function')\n",
|
|
"\n",
|
|
"plt.show()\n",
|
|
"\n",
|
|
"\"\"\"Plots a graph of the squashing function used by a rectified linear\n",
|
|
"unit\"\"\"\n",
|
|
"z = numpy.arange(-2, 2, .1)\n",
|
|
"zero = numpy.zeros(len(z))\n",
|
|
"y = numpy.max([zero, z], axis=0)\n",
|
|
"\n",
|
|
"fig = plt.figure()\n",
|
|
"ax = fig.add_subplot(111)\n",
|
|
"ax.plot(z, y)\n",
|
|
"ax.set_ylim([-2.0, 2.0])\n",
|
|
"ax.set_xlim([-2.0, 2.0])\n",
|
|
"ax.grid(True)\n",
|
|
"ax.set_xlabel('z')\n",
|
|
"ax.set_title('Rectified linear unit')\n",
|
|
"\n",
|
|
"plt.show()"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "code",
|
|
"execution_count": 2,
|
|
"metadata": {
|
|
"collapsed": false
|
|
},
|
|
"outputs": [],
|
|
"source": [
|
|
"from scipy import optimize\n",
|
|
"\n",
|
|
"class Neural_Network(object):\n",
|
|
" def __init__(self, Lambda=0): \n",
|
|
" #Define Hyperparameters\n",
|
|
" self.inputLayerSize = 2\n",
|
|
" self.outputLayerSize = 1\n",
|
|
" self.hiddenLayerSize = 3\n",
|
|
" \n",
|
|
" #Weights (parameters)\n",
|
|
" self.W1 = np.random.randn(self.inputLayerSize,self.hiddenLayerSize)\n",
|
|
" self.W2 = np.random.randn(self.hiddenLayerSize,self.outputLayerSize)\n",
|
|
" \n",
|
|
" #Regularization Parameter:\n",
|
|
" self.Lambda = Lambda\n",
|
|
" \n",
|
|
" def forward(self, X):\n",
|
|
" #Propogate inputs though network\n",
|
|
" self.z2 = np.dot(X, self.W1)\n",
|
|
" self.a2 = self.sigmoid(self.z2)\n",
|
|
" self.z3 = np.dot(self.a2, self.W2)\n",
|
|
" yHat = self.sigmoid(self.z3) \n",
|
|
" return yHat\n",
|
|
" \n",
|
|
" def sigmoid(self, z):\n",
|
|
" #Apply sigmoid activation function to scalar, vector, or matrix\n",
|
|
" return 1/(1+np.exp(-z))\n",
|
|
" \n",
|
|
" def sigmoidPrime(self,z):\n",
|
|
" #Gradient of sigmoid\n",
|
|
" return np.exp(-z)/((1+np.exp(-z))**2)\n",
|
|
" \n",
|
|
" def costFunction(self, X, y):\n",
|
|
" #Compute cost for given X,y, use weights already stored in class.\n",
|
|
" self.yHat = self.forward(X)\n",
|
|
" J = 0.5*sum((y-self.yHat)**2)/X.shape[0] + (self.Lambda/2)*(np.sum(self.W1**2)+np.sum(self.W2**2))\n",
|
|
" return J\n",
|
|
" \n",
|
|
" def costFunctionPrime(self, X, y):\n",
|
|
" #Compute derivative with respect to W and W2 for a given X and y:\n",
|
|
" self.yHat = self.forward(X)\n",
|
|
" \n",
|
|
" delta3 = np.multiply(-(y-self.yHat), self.sigmoidPrime(self.z3))\n",
|
|
" #Add gradient of regularization term:\n",
|
|
" dJdW2 = np.dot(self.a2.T, delta3)/X.shape[0] + self.Lambda*self.W2\n",
|
|
" \n",
|
|
" delta2 = np.dot(delta3, self.W2.T)*self.sigmoidPrime(self.z2)\n",
|
|
" #Add gradient of regularization term:\n",
|
|
" dJdW1 = np.dot(X.T, delta2)/X.shape[0] + self.Lambda*self.W1\n",
|
|
" \n",
|
|
" return dJdW1, dJdW2\n",
|
|
" \n",
|
|
" #Helper functions for interacting with other methods/classes\n",
|
|
" def getParams(self):\n",
|
|
" #Get W1 and W2 Rolled into vector:\n",
|
|
" params = np.concatenate((self.W1.ravel(), self.W2.ravel()))\n",
|
|
" return params\n",
|
|
" \n",
|
|
" def setParams(self, params):\n",
|
|
" #Set W1 and W2 using single parameter vector:\n",
|
|
" W1_start = 0\n",
|
|
" W1_end = self.hiddenLayerSize*self.inputLayerSize\n",
|
|
" self.W1 = np.reshape(params[W1_start:W1_end], \\\n",
|
|
" (self.inputLayerSize, self.hiddenLayerSize))\n",
|
|
" W2_end = W1_end + self.hiddenLayerSize*self.outputLayerSize\n",
|
|
" self.W2 = np.reshape(params[W1_end:W2_end], \\\n",
|
|
" (self.hiddenLayerSize, self.outputLayerSize))\n",
|
|
" \n",
|
|
" def computeGradients(self, X, y):\n",
|
|
" dJdW1, dJdW2 = self.costFunctionPrime(X, y)\n",
|
|
" return np.concatenate((dJdW1.ravel(), dJdW2.ravel()))\n",
|
|
" \n",
|
|
" \n",
|
|
"class trainer(object):\n",
|
|
" def __init__(self, N):\n",
|
|
" #Make Local reference to network:\n",
|
|
" self.N = N\n",
|
|
" \n",
|
|
" def callbackF(self, params):\n",
|
|
" self.N.setParams(params)\n",
|
|
" self.J.append(self.N.costFunction(self.X, self.y))\n",
|
|
" self.testJ.append(self.N.costFunction(self.testX, self.testY))\n",
|
|
" \n",
|
|
" def costFunctionWrapper(self, params, X, y):\n",
|
|
" self.N.setParams(params)\n",
|
|
" cost = self.N.costFunction(X, y)\n",
|
|
" grad = self.N.computeGradients(X,y)\n",
|
|
" return cost, grad\n",
|
|
" \n",
|
|
" def train(self, trainX, trainY, testX, testY):\n",
|
|
" #Make an internal variable for the callback function:\n",
|
|
" self.X = trainX\n",
|
|
" self.y = trainY\n",
|
|
" \n",
|
|
" self.testX = testX\n",
|
|
" self.testY = testY\n",
|
|
"\n",
|
|
" #Make empty list to store training costs:\n",
|
|
" self.J = []\n",
|
|
" self.testJ = []\n",
|
|
" \n",
|
|
" params0 = self.N.getParams()\n",
|
|
"\n",
|
|
" options = {'maxiter': 200, 'disp' : True}\n",
|
|
" _res = optimize.minimize(self.costFunctionWrapper, params0, jac=True, method='BFGS', \\\n",
|
|
" args=(trainX, trainY), options=options, callback=self.callbackF)\n",
|
|
"\n",
|
|
" self.N.setParams(_res.x)\n",
|
|
" self.optimizationResults = _res"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"metadata": {},
|
|
"source": [
|
|
"<!-- !split -->\n",
|
|
"## Two-layer Neural Network"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "code",
|
|
"execution_count": 3,
|
|
"metadata": {
|
|
"collapsed": false
|
|
},
|
|
"outputs": [],
|
|
"source": [
|
|
"import numpy as np\n",
|
|
"\n",
|
|
"#sigmoid\n",
|
|
"def nonlin(x, deriv=False):\n",
|
|
" if (deriv==True):\n",
|
|
" return x*(1-x)\n",
|
|
" return 1/(1+np.exp(-x))\n",
|
|
"\n",
|
|
"#input data\n",
|
|
"x=np.array([[0,0,1],[0,1,1],[1,0,1],[1,1,1]])\n",
|
|
"\n",
|
|
"#output data\n",
|
|
"y=np.array([0,1,1,0]).T\n",
|
|
"\n",
|
|
"#seed random numbers to make calculation\n",
|
|
"np.random.seed(1)\n",
|
|
"\n",
|
|
"#initialize weights with mean=0\n",
|
|
"syn0=2*np.random.random((3,4))-1\n",
|
|
"\n",
|
|
"for iter in range(10000):\n",
|
|
" #forward propogation\n",
|
|
" l0=x\n",
|
|
" l1=nonlin(np.dot(l0,syn0))\n",
|
|
" l1_error=y-l1\n",
|
|
" #multiply error by slope of sigmoid at values of l1\n",
|
|
" l1_delta=l1_error*nonlin(l1,True)\n",
|
|
" #update weights\n",
|
|
" syn0+=np.dot(l0.T, l1_delta)\n",
|
|
" \n",
|
|
"print(\"Output after training: \",l1 )"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "code",
|
|
"execution_count": 4,
|
|
"metadata": {
|
|
"collapsed": false
|
|
},
|
|
"outputs": [],
|
|
"source": [
|
|
"import numpy as np\n",
|
|
"import random\n",
|
|
"class Network(object):\n",
|
|
" \n",
|
|
" def _init_(self, sizes):\n",
|
|
" self.num_layers=len(sizes)\n",
|
|
" self.sizes=sizes\n",
|
|
" self.biases=[np.random.randn(y,1) for y in sizes[1:]]\n",
|
|
" self.weights=[np.random.randn(y,x) for x,y in zip(sizes[:-1], sizes[1:])]\n",
|
|
"\n",
|
|
"#sizes is the number of neurons in each layer\n",
|
|
"#for example, say n_1st_layer=3, n_2nd_layer=3, n_3rd_layer=1, then net=Network([3,3,1])\n",
|
|
"\n",
|
|
"#The biases and weights are initialized randomly, using Gaussian distributions of mean=0, stdev=1\n",
|
|
"#z is a vector (or a np.array)\n",
|
|
"\n",
|
|
" def feedforward(self,a):\n",
|
|
" #returns output w/ 'a' as an input\n",
|
|
" for b, w in zip(self.biases, self.weights):\n",
|
|
" a=sigmoid(np.dot(w,b)+b)\n",
|
|
" return a\n",
|
|
" \n",
|
|
"#Apply a Stochastic Gradient Descent (SGD) method:\n",
|
|
" def SGD(self, training_data, epochs, mini_batch_size, eta, test_data=None):\n",
|
|
" \"\"\"Trains network using batches incorporating SGD. The network will be evaluated against the\n",
|
|
" test data after each epoch, with partial progress being printed out (this is useful for tracking,\n",
|
|
" but slows the process.)\"\"\"\n",
|
|
" if test_data: n_test=len(test_data)\n",
|
|
" n=len(training_data)\n",
|
|
" for j in xrange(epochs):\n",
|
|
" random.shuffle(training_data)\n",
|
|
" mini_batches=[training_data[k:k+mini_batch_size] for k in xrange(o,n,mini_batch_size)]\n",
|
|
" for mini_batch in mini_batches:\n",
|
|
" self.update_mini_batch(mini_batch, eta)\n",
|
|
" if test_data:\n",
|
|
" print (\"Epoch {0}: {1}/{2}\".format(j, self.evaluate(test_data), n_test))\n",
|
|
" else:\n",
|
|
" print (\"Epoch {0} complete\".format(j))\n",
|
|
" \n",
|
|
" \n",
|
|
" def update_mini_batch(self, mini_batch, eta):\n",
|
|
" #updates w and b using backpropagation to a single mini batch. eta is the learning rate.\"\n",
|
|
" nabla_b=[np.zeros(b.shape) for b in self.biases]\n",
|
|
" nabla_w=[np.zeros(w.shape) for w in self.weights]\n",
|
|
" for x,y in mini_batch:\n",
|
|
" delta_nabla_b, delta_nabla_w=self.backprop(x,y)\n",
|
|
" nabla_b=[nb+dnb for nb, dnb in zip(nabla_b, delta_nabla_b)]\n",
|
|
" nabla_w=[nw+dnw for nw, dnw in zip(nabla_w, delta_nabla_w)]\n",
|
|
" self.weights=[w-(eta/len(mini_batch))*nw for w, nw in zip(self.weights, nabla_w)]\n",
|
|
" self.biases=[b-(eta/len(mini_batch))*nb for b, nb in zip(self.biases, nabla_b)]\n",
|
|
" \n",
|
|
" def backprop(self, x, y):\n",
|
|
" \"\"\"Return a tuple ``(nabla_b, nabla_w)`` representing the\n",
|
|
" gradient for the cost function C_x. ``nabla_b`` and\n",
|
|
" ``nabla_w`` are layer-by-layer lists of numpy arrays, similar\n",
|
|
" to ``self.biases`` and ``self.weights``.\"\"\"\n",
|
|
" nabla_b = [np.zeros(b.shape) for b in self.biases]\n",
|
|
" nabla_w = [np.zeros(w.shape) for w in self.weights]\n",
|
|
" # feedforward\n",
|
|
" activation = x\n",
|
|
" activations = [x] # list to store all the activations, layer by layer\n",
|
|
" zs = [] # list to store all the z vectors, layer by layer\n",
|
|
" for b, w in zip(self.biases, self.weights):\n",
|
|
" z = np.dot(w, activation)+b\n",
|
|
" zs.append(z)\n",
|
|
" activation = sigmoid(z)\n",
|
|
" activations.append(activation)\n",
|
|
" # backward pass\n",
|
|
" delta = self.cost_derivative(activations[-1], y) * \\\n",
|
|
" sigmoid_prime(zs[-1])\n",
|
|
" nabla_b[-1] = delta\n",
|
|
" nabla_w[-1] = np.dot(delta, activations[-2].transpose())\n",
|
|
" # Note that the variable l in the loop below is used a little\n",
|
|
" # differently to the notation in Chapter 2 of the book. Here,\n",
|
|
" # l = 1 means the last layer of neurons, l = 2 is the\n",
|
|
" # second-last layer, and so on. It's a renumbering of the\n",
|
|
" # scheme in the book, used here to take advantage of the fact\n",
|
|
" # that Python can use negative indices in lists.\n",
|
|
" for l in xrange(2, self.num_layers):\n",
|
|
" z = zs[-l]\n",
|
|
" sp = sigmoid_prime(z)\n",
|
|
" delta = np.dot(self.weights[-l+1].transpose(), delta) * sp\n",
|
|
" nabla_b[-l] = delta\n",
|
|
" nabla_w[-l] = np.dot(delta, activations[-l-1].transpose())\n",
|
|
" return (nabla_b, nabla_w)\n",
|
|
"\n",
|
|
" def evaluate(self, test_data):\n",
|
|
" \"\"\"Return the number of test inputs for which the neural\n",
|
|
" network outputs the correct result. Note that the neural\n",
|
|
" network's output is assumed to be the index of whichever\n",
|
|
" neuron in the final layer has the highest activation.\"\"\"\n",
|
|
" test_results = [(np.argmax(self.feedforward(x)), y)\n",
|
|
" for (x, y) in test_data]\n",
|
|
" return sum(int(x == y) for (x, y) in test_results)\n",
|
|
"\n",
|
|
" def cost_derivative(self, output_activations, y):\n",
|
|
" \"\"\"Return the vector of partial derivatives \\partial C_x /\n",
|
|
" \\partial a for the output activations.\"\"\"\n",
|
|
" return (output_activations-y)\n",
|
|
" \n",
|
|
" \n",
|
|
" \n",
|
|
"#Functions\n",
|
|
"def sigmoid(z):\n",
|
|
" return 1.0/(1.0+np.exp(-z))\n",
|
|
"\n",
|
|
"def sigmoid_prime(z):\n",
|
|
" return sigmoid(z)*(1-sigmoid(z))\n",
|
|
"\n",
|
|
"network=Network()"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "code",
|
|
"execution_count": 5,
|
|
"metadata": {
|
|
"collapsed": false
|
|
},
|
|
"outputs": [],
|
|
"source": [
|
|
"# %load neural-networks-and-deep-learning/src/mnist_loader.py\n",
|
|
"\"\"\"\n",
|
|
"mnist_loader\n",
|
|
"~~~~~~~~~~~~\n",
|
|
"\n",
|
|
"A library to load the MNIST image data. For details of the data\n",
|
|
"structures that are returned, see the doc strings for ``load_data``\n",
|
|
"and ``load_data_wrapper``. In practice, ``load_data_wrapper`` is the\n",
|
|
"function usually called by our neural network code.\n",
|
|
"\"\"\"\n",
|
|
"\n",
|
|
"#### Libraries\n",
|
|
"# Standard library\n",
|
|
"import pickle\n",
|
|
"import gzip\n",
|
|
"\n",
|
|
"# Third-party libraries\n",
|
|
"import numpy as np\n",
|
|
"\n",
|
|
"def load_data():\n",
|
|
" \"\"\"Return the MNIST data as a tuple containing the training data,\n",
|
|
" the validation data, and the test data.\n",
|
|
"\n",
|
|
" The ``training_data`` is returned as a tuple with two entries.\n",
|
|
" The first entry contains the actual training images. This is a\n",
|
|
" numpy ndarray with 50,000 entries. Each entry is, in turn, a\n",
|
|
" numpy ndarray with 784 values, representing the 28 * 28 = 784\n",
|
|
" pixels in a single MNIST image.\n",
|
|
"\n",
|
|
" The second entry in the ``training_data`` tuple is a numpy ndarray\n",
|
|
" containing 50,000 entries. Those entries are just the digit\n",
|
|
" values (0...9) for the corresponding images contained in the first\n",
|
|
" entry of the tuple.\n",
|
|
"\n",
|
|
" The ``validation_data`` and ``test_data`` are similar, except\n",
|
|
" each contains only 10,000 images.\n",
|
|
"\n",
|
|
" This is a nice data format, but for use in neural networks it's\n",
|
|
" helpful to modify the format of the ``training_data`` a little.\n",
|
|
" That's done in the wrapper function ``load_data_wrapper()``, see\n",
|
|
" below.\n",
|
|
" \"\"\"\n",
|
|
" f = gzip.open('../data/mnist.pkl.gz', 'rb')\n",
|
|
" training_data, validation_data, test_data = cPickle.load(f)\n",
|
|
" f.close()\n",
|
|
" return (training_data, validation_data, test_data)\n",
|
|
"\n",
|
|
"def load_data_wrapper():\n",
|
|
" \"\"\"Return a tuple containing ``(training_data, validation_data,\n",
|
|
" test_data)``. Based on ``load_data``, but the format is more\n",
|
|
" convenient for use in our implementation of neural networks.\n",
|
|
"\n",
|
|
" In particular, ``training_data`` is a list containing 50,000\n",
|
|
" 2-tuples ``(x, y)``. ``x`` is a 784-dimensional numpy.ndarray\n",
|
|
" containing the input image. ``y`` is a 10-dimensional\n",
|
|
" numpy.ndarray representing the unit vector corresponding to the\n",
|
|
" correct digit for ``x``.\n",
|
|
"\n",
|
|
" ``validation_data`` and ``test_data`` are lists containing 10,000\n",
|
|
" 2-tuples ``(x, y)``. In each case, ``x`` is a 784-dimensional\n",
|
|
" numpy.ndarry containing the input image, and ``y`` is the\n",
|
|
" corresponding classification, i.e., the digit values (integers)\n",
|
|
" corresponding to ``x``.\n",
|
|
"\n",
|
|
" Obviously, this means we're using slightly different formats for\n",
|
|
" the training data and the validation / test data. These formats\n",
|
|
" turn out to be the most convenient for use in our neural network\n",
|
|
" code.\"\"\"\n",
|
|
" tr_d, va_d, te_d = load_data()\n",
|
|
" training_inputs = [np.reshape(x, (784, 1)) for x in tr_d[0]]\n",
|
|
" training_results = [vectorized_result(y) for y in tr_d[1]]\n",
|
|
" training_data = zip(training_inputs, training_results)\n",
|
|
" validation_inputs = [np.reshape(x, (784, 1)) for x in va_d[0]]\n",
|
|
" validation_data = zip(validation_inputs, va_d[1])\n",
|
|
" test_inputs = [np.reshape(x, (784, 1)) for x in te_d[0]]\n",
|
|
" test_data = zip(test_inputs, te_d[1])\n",
|
|
" return (training_data, validation_data, test_data)\n",
|
|
"\n",
|
|
"def vectorized_result(j):\n",
|
|
" \"\"\"Return a 10-dimensional unit vector with a 1.0 in the jth\n",
|
|
" position and zeroes elsewhere. This is used to convert a digit\n",
|
|
" (0...9) into a corresponding desired output from the neural\n",
|
|
" network.\"\"\"\n",
|
|
" e = np.zeros((10, 1))\n",
|
|
" e[j] = 1.0\n",
|
|
" return e\n",
|
|
"\n",
|
|
"net=network.Network([784,30,30])\n",
|
|
"net.SGD(training_data,30,10,3,test_data=test_data)"
|
|
]
|
|
}
|
|
],
|
|
"metadata": {},
|
|
"nbformat": 4,
|
|
"nbformat_minor": 2
|
|
}
|