update of exercises week 43
This commit is contained in:
Binary file not shown.
Binary file not shown.
Binary file not shown.
@@ -2,8 +2,10 @@
|
|||||||
"cells": [
|
"cells": [
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "ae832182",
|
"id": "d5ebb4c0",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"<!-- HTML file automatically generated from DocOnce source (https://github.com/doconce/doconce/)\n",
|
"<!-- HTML file automatically generated from DocOnce source (https://github.com/doconce/doconce/)\n",
|
||||||
"doconce format html exercisesweek43.do.txt -->\n",
|
"doconce format html exercisesweek43.do.txt -->\n",
|
||||||
@@ -12,8 +14,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "77f844d9",
|
"id": "812b4e46",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"# Exercises weeks 43 and 44 \n",
|
"# Exercises weeks 43 and 44 \n",
|
||||||
"**October 23-27, 2023**\n",
|
"**October 23-27, 2023**\n",
|
||||||
@@ -25,8 +29,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "d5b983e3",
|
"id": "3230cd2f",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"# Overarching aims of the exercises weeks 43 and 44\n",
|
"# Overarching aims of the exercises weeks 43 and 44\n",
|
||||||
"\n",
|
"\n",
|
||||||
@@ -63,8 +69,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "7f6cc12f",
|
"id": "e3617d4e",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"## The AND and XOR Gates\n",
|
"## The AND and XOR Gates\n",
|
||||||
"\n",
|
"\n",
|
||||||
@@ -99,8 +107,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "c0b07e46",
|
"id": "54e1e7fc",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"## Representing the Data Sets\n",
|
"## Representing the Data Sets\n",
|
||||||
"\n",
|
"\n",
|
||||||
@@ -109,8 +119,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "e0a9840a",
|
"id": "6b3c15cb",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"$$\n",
|
"$$\n",
|
||||||
"\\boldsymbol{X}=\\begin{bmatrix} 0 & 0 \\\\\n",
|
"\\boldsymbol{X}=\\begin{bmatrix} 0 & 0 \\\\\n",
|
||||||
@@ -122,8 +134,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "ff95d379",
|
"id": "acc25271",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"while the vector of outputs is $\\boldsymbol{y}^T=[0,1,1,0]$ for the XOR gate, $\\boldsymbol{y}^T=[0,0,0,1]$ for the AND gate and $\\boldsymbol{y}^T=[0,1,1,1]$ for the OR gate.\n",
|
"while the vector of outputs is $\\boldsymbol{y}^T=[0,1,1,0]$ for the XOR gate, $\\boldsymbol{y}^T=[0,0,0,1]$ for the AND gate and $\\boldsymbol{y}^T=[0,1,1,1]$ for the OR gate.\n",
|
||||||
"\n",
|
"\n",
|
||||||
@@ -150,8 +164,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "36fa3466",
|
"id": "ffe0a840",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"## Setting up dimensionalities by hand\n",
|
"## Setting up dimensionalities by hand\n",
|
||||||
"\n",
|
"\n",
|
||||||
@@ -160,8 +176,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "e2e5e808",
|
"id": "46abf545",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"$$\n",
|
"$$\n",
|
||||||
"\\boldsymbol{W_h}=\\begin{bmatrix} 1 & 1 \\\\\n",
|
"\\boldsymbol{W_h}=\\begin{bmatrix} 1 & 1 \\\\\n",
|
||||||
@@ -171,16 +189,20 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "419e3904",
|
"id": "86e105cc",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"Multiplying $\\boldsymbol{X}$ and $\\boldsymbol{W}$ gives"
|
"Multiplying $\\boldsymbol{X}$ and $\\boldsymbol{W}$ gives"
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "88a9712a",
|
"id": "5e1d21f1",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"$$\n",
|
"$$\n",
|
||||||
"\\boldsymbol{X}{W}_h=\\begin{bmatrix} 0 & 0 \\\\\n",
|
"\\boldsymbol{X}{W}_h=\\begin{bmatrix} 0 & 0 \\\\\n",
|
||||||
@@ -192,16 +214,20 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "036eee14",
|
"id": "692c6cbf",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"Assume also that the bias vector for the hidden layer is"
|
"Assume also that the bias vector for the hidden layer is"
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "3c771918",
|
"id": "5d81e641",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"$$\n",
|
"$$\n",
|
||||||
"\\boldsymbol{b}_h=\\begin{bmatrix} 0 \\\\\n",
|
"\\boldsymbol{b}_h=\\begin{bmatrix} 0 \\\\\n",
|
||||||
@@ -211,16 +237,20 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "f5a4150d",
|
"id": "a42bdee4",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"Adding it gives us the input to the activation function of the hidden layer"
|
"Adding it gives us the input to the activation function of the hidden layer"
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "e4eb7753",
|
"id": "5116b854",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"$$\n",
|
"$$\n",
|
||||||
"\\boldsymbol{z}_h=\\boldsymbol{X}\\boldsymbol{W}_h+\\boldsymbol{b}_h=\\begin{bmatrix} 0 & -1 \\\\\n",
|
"\\boldsymbol{z}_h=\\boldsymbol{X}\\boldsymbol{W}_h+\\boldsymbol{b}_h=\\begin{bmatrix} 0 & -1 \\\\\n",
|
||||||
@@ -232,16 +262,20 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "b5442b32",
|
"id": "8dd26d09",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"Let us then assume that our activation function is the RELU function, which simply means that we take the max of $0$ and the elements of the input argument $\\boldsymbol{z}_h$, that is we have"
|
"Let us then assume that our activation function is the RELU function, which simply means that we take the max of $0$ and the elements of the input argument $\\boldsymbol{z}_h$, that is we have"
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "64d99f85",
|
"id": "ebf97c03",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"$$\n",
|
"$$\n",
|
||||||
"\\boldsymbol{a}_h=\\mathrm{RELU}(\\boldsymbol{z}_h=\\boldsymbol{X}\\boldsymbol{W}_h+\\boldsymbol{b}_h)=\\begin{bmatrix} 0 & 0 \\\\\n",
|
"\\boldsymbol{a}_h=\\mathrm{RELU}(\\boldsymbol{z}_h=\\boldsymbol{X}\\boldsymbol{W}_h+\\boldsymbol{b}_h)=\\begin{bmatrix} 0 & 0 \\\\\n",
|
||||||
@@ -253,16 +287,20 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "2aee4b6b",
|
"id": "262a4a3a",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"Assume also that the bias of the output layer is zero and that the weights of the output layer are"
|
"Assume also that the bias of the output layer is zero and that the weights of the output layer are"
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "884541eb",
|
"id": "f03ad8e7",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"$$\n",
|
"$$\n",
|
||||||
"\\boldsymbol{w}_o=\\begin{bmatrix} 1 \\\\\n",
|
"\\boldsymbol{w}_o=\\begin{bmatrix} 1 \\\\\n",
|
||||||
@@ -272,37 +310,46 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "3999a4d5",
|
"id": "36fe00c0",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"and multiplying with $\\boldsymbol{a}_h$ gives the output"
|
"and multiplying with $\\boldsymbol{a}_h$ gives the output"
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "a4896b4d",
|
"id": "80ab43ac",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"$$\n",
|
"$$\n",
|
||||||
"\\boldsymbol{a}_o=\\boldsymbol{w}_h^T\\begin{bmatrix} 0 & 0 \\\\\n",
|
"\\boldsymbol{a}_o=\\begin{bmatrix} 0 & 0 \\\\\n",
|
||||||
" 1 & 0 \\\\\n",
|
" 1 & 0 \\\\\n",
|
||||||
"\t\t 1 & 0 \\\\\n",
|
"\t\t 1 & 0 \\\\\n",
|
||||||
"\t\t 2 & 1 \\end{bmatrix}=\\begin{bmatrix} 0 \\\\ 1 \\\\ 1 \\\\0\\end{bmatrix},\n",
|
"\t\t 2 & 1 \\end{bmatrix}\\begin{bmatrix} 1 \\\\\n",
|
||||||
|
" -2\\end{bmatrix}=\\begin{bmatrix} 0 \\\\ 1 \\\\ 1 \\\\0\\end{bmatrix},\n",
|
||||||
"$$"
|
"$$"
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "2c74e7c0",
|
"id": "86bcfe49",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"the wanted result."
|
"the wanted result. Pay attention to the dimensionalities as well."
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "249b9804",
|
"id": "f19a899e",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"## Setting up the Neural Network\n",
|
"## Setting up the Neural Network\n",
|
||||||
"\n",
|
"\n",
|
||||||
@@ -312,8 +359,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 1,
|
"execution_count": 1,
|
||||||
"id": "afdd3100",
|
"id": "901ddca7",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [],
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"%matplotlib inline\n",
|
"%matplotlib inline\n",
|
||||||
@@ -379,16 +429,20 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "fb383164",
|
"id": "53f22266",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"Not an impressive result, but this was our first forward pass with randomly assigned weights. Let us now add the full network with the back-propagation algorithm discussed above."
|
"Not an impressive result, but this was our first forward pass with randomly assigned weights. Let us now add the full network with the back-propagation algorithm discussed above."
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "0d26394c",
|
"id": "5f665ac6",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"## The Code using Scikit-Learn"
|
"## The Code using Scikit-Learn"
|
||||||
]
|
]
|
||||||
@@ -396,8 +450,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 2,
|
"execution_count": 2,
|
||||||
"id": "befc368c",
|
"id": "cb396cda",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [],
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"# import necessary packages\n",
|
"# import necessary packages\n",
|
||||||
@@ -458,8 +515,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "03895a9f",
|
"id": "c5978471",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"## Building a neural network code\n",
|
"## Building a neural network code\n",
|
||||||
"\n",
|
"\n",
|
||||||
@@ -475,8 +534,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "6226915b",
|
"id": "7b82dc46",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"### Learning rate methods\n",
|
"### Learning rate methods\n",
|
||||||
"\n",
|
"\n",
|
||||||
@@ -495,8 +556,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 3,
|
"execution_count": 3,
|
||||||
"id": "b6d5ea73",
|
"id": "de4dedfb",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [],
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"import autograd.numpy as np\n",
|
"import autograd.numpy as np\n",
|
||||||
@@ -633,8 +697,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "edb42bef",
|
"id": "bb621fce",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"### Usage of the above learning rate schedulers\n",
|
"### Usage of the above learning rate schedulers\n",
|
||||||
"\n",
|
"\n",
|
||||||
@@ -647,8 +713,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 4,
|
"execution_count": 4,
|
||||||
"id": "108c1209",
|
"id": "34ddb829",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [],
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"momentum_scheduler = Momentum(eta=1e-3, momentum=0.9)\n",
|
"momentum_scheduler = Momentum(eta=1e-3, momentum=0.9)\n",
|
||||||
@@ -657,8 +726,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "d1c423f0",
|
"id": "e43498d4",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"Here is a small example for how a segment of code using schedulers\n",
|
"Here is a small example for how a segment of code using schedulers\n",
|
||||||
"could look. Switching out the schedulers is simple."
|
"could look. Switching out the schedulers is simple."
|
||||||
@@ -667,8 +738,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 5,
|
"execution_count": 5,
|
||||||
"id": "63c8c749",
|
"id": "f05b9625",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [],
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"weights = np.ones((3,3))\n",
|
"weights = np.ones((3,3))\n",
|
||||||
@@ -686,8 +760,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "7cb80992",
|
"id": "19cf9841",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"### Cost functions\n",
|
"### Cost functions\n",
|
||||||
"\n",
|
"\n",
|
||||||
@@ -700,8 +776,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 6,
|
"execution_count": 6,
|
||||||
"id": "a181dacb",
|
"id": "993efa8d",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [],
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"import autograd.numpy as np\n",
|
"import autograd.numpy as np\n",
|
||||||
@@ -735,8 +814,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "26a7fe37",
|
"id": "011c734c",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"Below we give a short example of how these cost function may be used\n",
|
"Below we give a short example of how these cost function may be used\n",
|
||||||
"to obtain results if you wish to test them out on your own using\n",
|
"to obtain results if you wish to test them out on your own using\n",
|
||||||
@@ -746,8 +827,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 7,
|
"execution_count": 7,
|
||||||
"id": "5894bd31",
|
"id": "9e1b97f5",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [],
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"from autograd import grad\n",
|
"from autograd import grad\n",
|
||||||
@@ -764,8 +848,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "71141fab",
|
"id": "4acd87b2",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"### Activation functions\n",
|
"### Activation functions\n",
|
||||||
"\n",
|
"\n",
|
||||||
@@ -778,8 +864,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 8,
|
"execution_count": 8,
|
||||||
"id": "433e0886",
|
"id": "befa86ae",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [],
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"import autograd.numpy as np\n",
|
"import autograd.numpy as np\n",
|
||||||
@@ -833,8 +922,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "14c9d538",
|
"id": "c20aa75e",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"Below follows a short demonstration of how to use an activation\n",
|
"Below follows a short demonstration of how to use an activation\n",
|
||||||
"function. The derivative of the activation function will be important\n",
|
"function. The derivative of the activation function will be important\n",
|
||||||
@@ -846,8 +937,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 9,
|
"execution_count": 9,
|
||||||
"id": "4c3a44ef",
|
"id": "e209f9e5",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [],
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"z = np.array([[4, 5, 6]]).T\n",
|
"z = np.array([[4, 5, 6]]).T\n",
|
||||||
@@ -864,8 +958,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "aa7f604e",
|
"id": "524b409b",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"### The Neural Network\n",
|
"### The Neural Network\n",
|
||||||
"\n",
|
"\n",
|
||||||
@@ -886,8 +982,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 10,
|
"execution_count": 10,
|
||||||
"id": "585a227b",
|
"id": "083116d3",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [],
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"import math\n",
|
"import math\n",
|
||||||
@@ -1355,8 +1454,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "7de704af",
|
"id": "e4ba55d0",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"Before we make a model, we will quickly generate a dataset we can use\n",
|
"Before we make a model, we will quickly generate a dataset we can use\n",
|
||||||
"for our linear regression problem as shown below"
|
"for our linear regression problem as shown below"
|
||||||
@@ -1365,8 +1466,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 11,
|
"execution_count": 11,
|
||||||
"id": "212b0a35",
|
"id": "c72a22c6",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [],
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"import autograd.numpy as np\n",
|
"import autograd.numpy as np\n",
|
||||||
@@ -1406,8 +1510,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "1ff4ec7b",
|
"id": "25b63b47",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"Now that we have our dataset ready for the regression, we can create\n",
|
"Now that we have our dataset ready for the regression, we can create\n",
|
||||||
"our regressor. Note that with the seed parameter, we can make sure our\n",
|
"our regressor. Note that with the seed parameter, we can make sure our\n",
|
||||||
@@ -1420,8 +1526,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 12,
|
"execution_count": 12,
|
||||||
"id": "57d8eca3",
|
"id": "b6b4e461",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [],
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"input_nodes = X_train.shape[1]\n",
|
"input_nodes = X_train.shape[1]\n",
|
||||||
@@ -1432,8 +1541,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "f1977bae",
|
"id": "d7860b74",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"We then fit our model with our training data using the scheduler of our choice."
|
"We then fit our model with our training data using the scheduler of our choice."
|
||||||
]
|
]
|
||||||
@@ -1441,8 +1552,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 13,
|
"execution_count": 13,
|
||||||
"id": "31ae2a1e",
|
"id": "00522c73",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [],
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"linear_regression.reset_weights() # reset weights such that previous runs or reruns don't affect the weights\n",
|
"linear_regression.reset_weights() # reset weights such that previous runs or reruns don't affect the weights\n",
|
||||||
@@ -1453,8 +1567,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "96a64da4",
|
"id": "a57bfb12",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"Due to the progress bar we can see the MSE (train_error) throughout\n",
|
"Due to the progress bar we can see the MSE (train_error) throughout\n",
|
||||||
"the FFNN's training. Note that the fit() function has some optional\n",
|
"the FFNN's training. Note that the fit() function has some optional\n",
|
||||||
@@ -1467,8 +1583,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 14,
|
"execution_count": 14,
|
||||||
"id": "246499fc",
|
"id": "35259e41",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [],
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"linear_regression.reset_weights() # reset weights such that previous runs or reruns don't affect the weights\n",
|
"linear_regression.reset_weights() # reset weights such that previous runs or reruns don't affect the weights\n",
|
||||||
@@ -1478,8 +1597,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "61d11e17",
|
"id": "4a403370",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"We see that given more epochs to train on, the regressor reaches a lower MSE.\n",
|
"We see that given more epochs to train on, the regressor reaches a lower MSE.\n",
|
||||||
"\n",
|
"\n",
|
||||||
@@ -1491,8 +1612,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 15,
|
"execution_count": 15,
|
||||||
"id": "b5b91899",
|
"id": "6c791130",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [],
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"from sklearn.datasets import load_breast_cancer\n",
|
"from sklearn.datasets import load_breast_cancer\n",
|
||||||
@@ -1514,8 +1638,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 16,
|
"execution_count": 16,
|
||||||
"id": "4bd01b3d",
|
"id": "0ac0258d",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [],
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"input_nodes = X_train.shape[1]\n",
|
"input_nodes = X_train.shape[1]\n",
|
||||||
@@ -1526,8 +1653,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "3a138d66",
|
"id": "a9c74baa",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"We will now make use of our validation data by passing it into our fit function as a keyword argument"
|
"We will now make use of our validation data by passing it into our fit function as a keyword argument"
|
||||||
]
|
]
|
||||||
@@ -1535,8 +1664,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 17,
|
"execution_count": 17,
|
||||||
"id": "faf1534e",
|
"id": "e5021bbf",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [],
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"logistic_regression.reset_weights() # reset weights such that previous runs or reruns don't affect the weights\n",
|
"logistic_regression.reset_weights() # reset weights such that previous runs or reruns don't affect the weights\n",
|
||||||
@@ -1547,8 +1679,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "33d2d346",
|
"id": "ae0148bb",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"Finally, we will create a neural network with 2 hidden layers with activation functions."
|
"Finally, we will create a neural network with 2 hidden layers with activation functions."
|
||||||
]
|
]
|
||||||
@@ -1556,8 +1690,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 18,
|
"execution_count": 18,
|
||||||
"id": "a8f0a61c",
|
"id": "bb7e4340",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [],
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"input_nodes = X_train.shape[1]\n",
|
"input_nodes = X_train.shape[1]\n",
|
||||||
@@ -1573,8 +1710,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 19,
|
"execution_count": 19,
|
||||||
"id": "d651702d",
|
"id": "d0655afb",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [],
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"neural_network.reset_weights() # reset weights such that previous runs or reruns don't affect the weights\n",
|
"neural_network.reset_weights() # reset weights such that previous runs or reruns don't affect the weights\n",
|
||||||
@@ -1585,8 +1725,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "200b64c0",
|
"id": "dc8a90af",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"### Multiclass classification\n",
|
"### Multiclass classification\n",
|
||||||
"\n",
|
"\n",
|
||||||
@@ -1598,8 +1740,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 20,
|
"execution_count": 20,
|
||||||
"id": "98b8d59b",
|
"id": "9305c08a",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [],
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"from sklearn.datasets import load_digits\n",
|
"from sklearn.datasets import load_digits\n",
|
||||||
@@ -1632,8 +1777,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "cadd7e47",
|
"id": "8e208051",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"## Testing the XOR gate and other gates\n",
|
"## Testing the XOR gate and other gates\n",
|
||||||
"\n",
|
"\n",
|
||||||
@@ -1643,8 +1790,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 21,
|
"execution_count": 21,
|
||||||
"id": "a40e5cfe",
|
"id": "db016d41",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [],
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"X = np.array([ [0, 0], [0, 1], [1, 0],[1, 1]],dtype=np.float64)\n",
|
"X = np.array([ [0, 0], [0, 1], [1, 0],[1, 1]],dtype=np.float64)\n",
|
||||||
@@ -1663,32 +1813,16 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "b635db73",
|
"id": "d1384449",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"Not bad, but the results depend strongly on the learning reate. Try different learning rates."
|
"Not bad, but the results depend strongly on the learning reate. Try different learning rates."
|
||||||
]
|
]
|
||||||
}
|
}
|
||||||
],
|
],
|
||||||
"metadata": {
|
"metadata": {},
|
||||||
"kernelspec": {
|
|
||||||
"display_name": "Python 3 (ipykernel)",
|
|
||||||
"language": "python",
|
|
||||||
"name": "python3"
|
|
||||||
},
|
|
||||||
"language_info": {
|
|
||||||
"codemirror_mode": {
|
|
||||||
"name": "ipython",
|
|
||||||
"version": 3
|
|
||||||
},
|
|
||||||
"file_extension": ".py",
|
|
||||||
"mimetype": "text/x-python",
|
|
||||||
"name": "python",
|
|
||||||
"nbconvert_exporter": "python",
|
|
||||||
"pygments_lexer": "ipython3",
|
|
||||||
"version": "3.9.10"
|
|
||||||
}
|
|
||||||
},
|
|
||||||
"nbformat": 4,
|
"nbformat": 4,
|
||||||
"nbformat_minor": 5
|
"nbformat_minor": 5
|
||||||
}
|
}
|
||||||
|
|||||||
File diff suppressed because it is too large
Load Diff
@@ -750,12 +750,13 @@ inputs <span class="math notranslate nohighlight">\(x_1\)</span> and <span class
|
|||||||
<p>and multiplying with <span class="math notranslate nohighlight">\(\boldsymbol{a}_h\)</span> gives the output</p>
|
<p>and multiplying with <span class="math notranslate nohighlight">\(\boldsymbol{a}_h\)</span> gives the output</p>
|
||||||
<div class="math notranslate nohighlight">
|
<div class="math notranslate nohighlight">
|
||||||
\[\begin{split}
|
\[\begin{split}
|
||||||
\boldsymbol{a}_o=\boldsymbol{w}_h^T\begin{bmatrix} 0 & 0 \\
|
\boldsymbol{a}_o=\begin{bmatrix} 0 & 0 \\
|
||||||
1 & 0 \\
|
1 & 0 \\
|
||||||
1 & 0 \\
|
1 & 0 \\
|
||||||
2 & 1 \end{bmatrix}=\begin{bmatrix} 0 \\ 1 \\ 1 \\0\end{bmatrix},
|
2 & 1 \end{bmatrix}\begin{bmatrix} 1 \\
|
||||||
|
-2\end{bmatrix}=\begin{bmatrix} 0 \\ 1 \\ 1 \\0\end{bmatrix},
|
||||||
\end{split}\]</div>
|
\end{split}\]</div>
|
||||||
<p>the wanted result.</p>
|
<p>the wanted result. Pay attention to the dimensionalities as well.</p>
|
||||||
</div>
|
</div>
|
||||||
<div class="section" id="setting-up-the-neural-network">
|
<div class="section" id="setting-up-the-neural-network">
|
||||||
<h2>Setting up the Neural Network<a class="headerlink" href="#setting-up-the-neural-network" title="Permalink to this headline">¶</a></h2>
|
<h2>Setting up the Neural Network<a class="headerlink" href="#setting-up-the-neural-network" title="Permalink to this headline">¶</a></h2>
|
||||||
@@ -8456,8 +8457,9 @@ case.</p>
|
|||||||
</div>
|
</div>
|
||||||
<div class="cell_output docutils container">
|
<div class="cell_output docutils container">
|
||||||
<div class="output stream highlight-myst-ansi notranslate"><div class="highlight"><pre><span></span>Adam: Eta=0.0001, Lambda=0
|
<div class="output stream highlight-myst-ansi notranslate"><div class="highlight"><pre><span></span>Adam: Eta=0.0001, Lambda=0
|
||||||
|
</pre></div>
|
||||||
[----------------------------------------] 0.000% | train_error: 11.3 | train_acc: 0.453 | val_error: 10.4 | val_acc: 0.497
|
</div>
|
||||||
|
<div class="output stream highlight-myst-ansi notranslate"><div class="highlight"><pre><span></span> [----------------------------------------] 0.000% | train_error: 11.3 | train_acc: 0.453 | val_error: 10.4 | val_acc: 0.497
|
||||||
</pre></div>
|
</pre></div>
|
||||||
</div>
|
</div>
|
||||||
<div class="output stream highlight-myst-ansi notranslate"><div class="highlight"><pre><span></span> [----------------------------------------] 0.1000% | train_error: 11.4 | train_acc: 0.448 | val_error: 10.4 | val_acc: 0.497
|
<div class="output stream highlight-myst-ansi notranslate"><div class="highlight"><pre><span></span> [----------------------------------------] 0.1000% | train_error: 11.4 | train_acc: 0.448 | val_error: 10.4 | val_acc: 0.497
|
||||||
|
|||||||
File diff suppressed because one or more lines are too long
@@ -1689,12 +1689,13 @@ inputs <span class="math notranslate nohighlight">\(x_1\)</span> and <span class
|
|||||||
<p>and multiplying with <span class="math notranslate nohighlight">\(\boldsymbol{a}_h\)</span> gives the output</p>
|
<p>and multiplying with <span class="math notranslate nohighlight">\(\boldsymbol{a}_h\)</span> gives the output</p>
|
||||||
<div class="math notranslate nohighlight">
|
<div class="math notranslate nohighlight">
|
||||||
\[\begin{split}
|
\[\begin{split}
|
||||||
\boldsymbol{a}_o=\boldsymbol{w}_h^T\begin{bmatrix} 0 & 0 \\
|
\boldsymbol{a}_o=\begin{bmatrix} 0 & 0 \\
|
||||||
1 & 0 \\
|
1 & 0 \\
|
||||||
1 & 0 \\
|
1 & 0 \\
|
||||||
2 & 1 \end{bmatrix}=\begin{bmatrix} 0 \\ 1 \\ 1 \\0\end{bmatrix},
|
2 & 1 \end{bmatrix}\begin{bmatrix} 1 \\
|
||||||
|
-2\end{bmatrix}=\begin{bmatrix} 0 \\ 1 \\ 1 \\0\end{bmatrix},
|
||||||
\end{split}\]</div>
|
\end{split}\]</div>
|
||||||
<p>the wanted result.</p>
|
<p>the wanted result. Pay attention to the dimensionalities as well.</p>
|
||||||
</div>
|
</div>
|
||||||
<div class="section" id="setting-up-the-neural-network">
|
<div class="section" id="setting-up-the-neural-network">
|
||||||
<h2>Setting up the Neural Network<a class="headerlink" href="#setting-up-the-neural-network" title="Permalink to this headline">¶</a></h2>
|
<h2>Setting up the Neural Network<a class="headerlink" href="#setting-up-the-neural-network" title="Permalink to this headline">¶</a></h2>
|
||||||
@@ -6362,6 +6363,9 @@ case.</p>
|
|||||||
<div class="output stream highlight-myst-ansi notranslate"><div class="highlight"><pre><span></span>Adam: Eta=0.001, Lambda=0
|
<div class="output stream highlight-myst-ansi notranslate"><div class="highlight"><pre><span></span>Adam: Eta=0.001, Lambda=0
|
||||||
</pre></div>
|
</pre></div>
|
||||||
</div>
|
</div>
|
||||||
|
<div class="output stream highlight-myst-ansi notranslate"><div class="highlight"><pre><span></span>
|
||||||
|
</pre></div>
|
||||||
|
</div>
|
||||||
<div class="output stream highlight-myst-ansi notranslate"><div class="highlight"><pre><span></span> [----------------------------------------] 0.000% | train_error: 12.9 | train_acc: 0.376 | val_error: 12.6 | val_acc: 0.392
|
<div class="output stream highlight-myst-ansi notranslate"><div class="highlight"><pre><span></span> [----------------------------------------] 0.000% | train_error: 12.9 | train_acc: 0.376 | val_error: 12.6 | val_acc: 0.392
|
||||||
</pre></div>
|
</pre></div>
|
||||||
</div>
|
</div>
|
||||||
@@ -18823,9 +18827,7 @@ This is then passed through the activation:</p>
|
|||||||
</div>
|
</div>
|
||||||
<div class="cell_output docutils container">
|
<div class="cell_output docutils container">
|
||||||
<div class="output stream highlight-myst-ansi notranslate"><div class="highlight"><pre><span></span>probabilities = (n_inputs, n_categories) = (1437, 10)
|
<div class="output stream highlight-myst-ansi notranslate"><div class="highlight"><pre><span></span>probabilities = (n_inputs, n_categories) = (1437, 10)
|
||||||
</pre></div>
|
probability that image 0 is in category 0,1,2,...,9 =
|
||||||
</div>
|
|
||||||
<div class="output stream highlight-myst-ansi notranslate"><div class="highlight"><pre><span></span>probability that image 0 is in category 0,1,2,...,9 =
|
|
||||||
[5.41511965e-04 2.17174962e-03 8.84355903e-03 1.44970586e-03
|
[5.41511965e-04 2.17174962e-03 8.84355903e-03 1.44970586e-03
|
||||||
1.10378326e-04 5.08318298e-09 2.03256632e-04 1.92507116e-03
|
1.10378326e-04 5.08318298e-09 2.03256632e-04 1.92507116e-03
|
||||||
9.84443254e-01 3.11507992e-04]
|
9.84443254e-01 3.11507992e-04]
|
||||||
|
|||||||
@@ -2,8 +2,10 @@
|
|||||||
"cells": [
|
"cells": [
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "ae832182",
|
"id": "d5ebb4c0",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"<!-- HTML file automatically generated from DocOnce source (https://github.com/doconce/doconce/)\n",
|
"<!-- HTML file automatically generated from DocOnce source (https://github.com/doconce/doconce/)\n",
|
||||||
"doconce format html exercisesweek43.do.txt -->\n",
|
"doconce format html exercisesweek43.do.txt -->\n",
|
||||||
@@ -12,8 +14,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "77f844d9",
|
"id": "812b4e46",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"# Exercises weeks 43 and 44 \n",
|
"# Exercises weeks 43 and 44 \n",
|
||||||
"**October 23-27, 2023**\n",
|
"**October 23-27, 2023**\n",
|
||||||
@@ -25,8 +29,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "d5b983e3",
|
"id": "3230cd2f",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"# Overarching aims of the exercises weeks 43 and 44\n",
|
"# Overarching aims of the exercises weeks 43 and 44\n",
|
||||||
"\n",
|
"\n",
|
||||||
@@ -63,8 +69,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "7f6cc12f",
|
"id": "e3617d4e",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"## The AND and XOR Gates\n",
|
"## The AND and XOR Gates\n",
|
||||||
"\n",
|
"\n",
|
||||||
@@ -99,8 +107,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "c0b07e46",
|
"id": "54e1e7fc",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"## Representing the Data Sets\n",
|
"## Representing the Data Sets\n",
|
||||||
"\n",
|
"\n",
|
||||||
@@ -109,8 +119,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "e0a9840a",
|
"id": "6b3c15cb",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"$$\n",
|
"$$\n",
|
||||||
"\\boldsymbol{X}=\\begin{bmatrix} 0 & 0 \\\\\n",
|
"\\boldsymbol{X}=\\begin{bmatrix} 0 & 0 \\\\\n",
|
||||||
@@ -122,8 +134,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "ff95d379",
|
"id": "acc25271",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"while the vector of outputs is $\\boldsymbol{y}^T=[0,1,1,0]$ for the XOR gate, $\\boldsymbol{y}^T=[0,0,0,1]$ for the AND gate and $\\boldsymbol{y}^T=[0,1,1,1]$ for the OR gate.\n",
|
"while the vector of outputs is $\\boldsymbol{y}^T=[0,1,1,0]$ for the XOR gate, $\\boldsymbol{y}^T=[0,0,0,1]$ for the AND gate and $\\boldsymbol{y}^T=[0,1,1,1]$ for the OR gate.\n",
|
||||||
"\n",
|
"\n",
|
||||||
@@ -150,8 +164,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "36fa3466",
|
"id": "ffe0a840",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"## Setting up dimensionalities by hand\n",
|
"## Setting up dimensionalities by hand\n",
|
||||||
"\n",
|
"\n",
|
||||||
@@ -160,8 +176,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "e2e5e808",
|
"id": "46abf545",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"$$\n",
|
"$$\n",
|
||||||
"\\boldsymbol{W_h}=\\begin{bmatrix} 1 & 1 \\\\\n",
|
"\\boldsymbol{W_h}=\\begin{bmatrix} 1 & 1 \\\\\n",
|
||||||
@@ -171,16 +189,20 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "419e3904",
|
"id": "86e105cc",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"Multiplying $\\boldsymbol{X}$ and $\\boldsymbol{W}$ gives"
|
"Multiplying $\\boldsymbol{X}$ and $\\boldsymbol{W}$ gives"
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "88a9712a",
|
"id": "5e1d21f1",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"$$\n",
|
"$$\n",
|
||||||
"\\boldsymbol{X}{W}_h=\\begin{bmatrix} 0 & 0 \\\\\n",
|
"\\boldsymbol{X}{W}_h=\\begin{bmatrix} 0 & 0 \\\\\n",
|
||||||
@@ -192,16 +214,20 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "036eee14",
|
"id": "692c6cbf",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"Assume also that the bias vector for the hidden layer is"
|
"Assume also that the bias vector for the hidden layer is"
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "3c771918",
|
"id": "5d81e641",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"$$\n",
|
"$$\n",
|
||||||
"\\boldsymbol{b}_h=\\begin{bmatrix} 0 \\\\\n",
|
"\\boldsymbol{b}_h=\\begin{bmatrix} 0 \\\\\n",
|
||||||
@@ -211,16 +237,20 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "f5a4150d",
|
"id": "a42bdee4",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"Adding it gives us the input to the activation function of the hidden layer"
|
"Adding it gives us the input to the activation function of the hidden layer"
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "e4eb7753",
|
"id": "5116b854",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"$$\n",
|
"$$\n",
|
||||||
"\\boldsymbol{z}_h=\\boldsymbol{X}\\boldsymbol{W}_h+\\boldsymbol{b}_h=\\begin{bmatrix} 0 & -1 \\\\\n",
|
"\\boldsymbol{z}_h=\\boldsymbol{X}\\boldsymbol{W}_h+\\boldsymbol{b}_h=\\begin{bmatrix} 0 & -1 \\\\\n",
|
||||||
@@ -232,16 +262,20 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "b5442b32",
|
"id": "8dd26d09",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"Let us then assume that our activation function is the RELU function, which simply means that we take the max of $0$ and the elements of the input argument $\\boldsymbol{z}_h$, that is we have"
|
"Let us then assume that our activation function is the RELU function, which simply means that we take the max of $0$ and the elements of the input argument $\\boldsymbol{z}_h$, that is we have"
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "64d99f85",
|
"id": "ebf97c03",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"$$\n",
|
"$$\n",
|
||||||
"\\boldsymbol{a}_h=\\mathrm{RELU}(\\boldsymbol{z}_h=\\boldsymbol{X}\\boldsymbol{W}_h+\\boldsymbol{b}_h)=\\begin{bmatrix} 0 & 0 \\\\\n",
|
"\\boldsymbol{a}_h=\\mathrm{RELU}(\\boldsymbol{z}_h=\\boldsymbol{X}\\boldsymbol{W}_h+\\boldsymbol{b}_h)=\\begin{bmatrix} 0 & 0 \\\\\n",
|
||||||
@@ -253,16 +287,20 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "2aee4b6b",
|
"id": "262a4a3a",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"Assume also that the bias of the output layer is zero and that the weights of the output layer are"
|
"Assume also that the bias of the output layer is zero and that the weights of the output layer are"
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "884541eb",
|
"id": "f03ad8e7",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"$$\n",
|
"$$\n",
|
||||||
"\\boldsymbol{w}_o=\\begin{bmatrix} 1 \\\\\n",
|
"\\boldsymbol{w}_o=\\begin{bmatrix} 1 \\\\\n",
|
||||||
@@ -272,37 +310,46 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "3999a4d5",
|
"id": "36fe00c0",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"and multiplying with $\\boldsymbol{a}_h$ gives the output"
|
"and multiplying with $\\boldsymbol{a}_h$ gives the output"
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "a4896b4d",
|
"id": "80ab43ac",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"$$\n",
|
"$$\n",
|
||||||
"\\boldsymbol{a}_o=\\boldsymbol{w}_h^T\\begin{bmatrix} 0 & 0 \\\\\n",
|
"\\boldsymbol{a}_o=\\begin{bmatrix} 0 & 0 \\\\\n",
|
||||||
" 1 & 0 \\\\\n",
|
" 1 & 0 \\\\\n",
|
||||||
"\t\t 1 & 0 \\\\\n",
|
"\t\t 1 & 0 \\\\\n",
|
||||||
"\t\t 2 & 1 \\end{bmatrix}=\\begin{bmatrix} 0 \\\\ 1 \\\\ 1 \\\\0\\end{bmatrix},\n",
|
"\t\t 2 & 1 \\end{bmatrix}\\begin{bmatrix} 1 \\\\\n",
|
||||||
|
" -2\\end{bmatrix}=\\begin{bmatrix} 0 \\\\ 1 \\\\ 1 \\\\0\\end{bmatrix},\n",
|
||||||
"$$"
|
"$$"
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "2c74e7c0",
|
"id": "86bcfe49",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"the wanted result."
|
"the wanted result. Pay attention to the dimensionalities as well."
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "249b9804",
|
"id": "f19a899e",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"## Setting up the Neural Network\n",
|
"## Setting up the Neural Network\n",
|
||||||
"\n",
|
"\n",
|
||||||
@@ -312,8 +359,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 1,
|
"execution_count": 1,
|
||||||
"id": "afdd3100",
|
"id": "901ddca7",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [
|
"outputs": [
|
||||||
{
|
{
|
||||||
"name": "stdout",
|
"name": "stdout",
|
||||||
@@ -390,16 +440,20 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "fb383164",
|
"id": "53f22266",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"Not an impressive result, but this was our first forward pass with randomly assigned weights. Let us now add the full network with the back-propagation algorithm discussed above."
|
"Not an impressive result, but this was our first forward pass with randomly assigned weights. Let us now add the full network with the back-propagation algorithm discussed above."
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "0d26394c",
|
"id": "5f665ac6",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"## The Code using Scikit-Learn"
|
"## The Code using Scikit-Learn"
|
||||||
]
|
]
|
||||||
@@ -407,8 +461,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 2,
|
"execution_count": 2,
|
||||||
"id": "befc368c",
|
"id": "cb396cda",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [
|
"outputs": [
|
||||||
{
|
{
|
||||||
"name": "stdout",
|
"name": "stdout",
|
||||||
@@ -712,8 +769,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "03895a9f",
|
"id": "c5978471",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"## Building a neural network code\n",
|
"## Building a neural network code\n",
|
||||||
"\n",
|
"\n",
|
||||||
@@ -729,8 +788,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "6226915b",
|
"id": "7b82dc46",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"### Learning rate methods\n",
|
"### Learning rate methods\n",
|
||||||
"\n",
|
"\n",
|
||||||
@@ -749,8 +810,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 3,
|
"execution_count": 3,
|
||||||
"id": "b6d5ea73",
|
"id": "de4dedfb",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [],
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"import autograd.numpy as np\n",
|
"import autograd.numpy as np\n",
|
||||||
@@ -887,8 +951,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "edb42bef",
|
"id": "bb621fce",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"### Usage of the above learning rate schedulers\n",
|
"### Usage of the above learning rate schedulers\n",
|
||||||
"\n",
|
"\n",
|
||||||
@@ -901,8 +967,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 4,
|
"execution_count": 4,
|
||||||
"id": "108c1209",
|
"id": "34ddb829",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [],
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"momentum_scheduler = Momentum(eta=1e-3, momentum=0.9)\n",
|
"momentum_scheduler = Momentum(eta=1e-3, momentum=0.9)\n",
|
||||||
@@ -911,8 +980,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "d1c423f0",
|
"id": "e43498d4",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"Here is a small example for how a segment of code using schedulers\n",
|
"Here is a small example for how a segment of code using schedulers\n",
|
||||||
"could look. Switching out the schedulers is simple."
|
"could look. Switching out the schedulers is simple."
|
||||||
@@ -921,8 +992,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 5,
|
"execution_count": 5,
|
||||||
"id": "63c8c749",
|
"id": "f05b9625",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [
|
"outputs": [
|
||||||
{
|
{
|
||||||
"name": "stdout",
|
"name": "stdout",
|
||||||
@@ -956,8 +1030,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "7cb80992",
|
"id": "19cf9841",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"### Cost functions\n",
|
"### Cost functions\n",
|
||||||
"\n",
|
"\n",
|
||||||
@@ -970,8 +1046,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 6,
|
"execution_count": 6,
|
||||||
"id": "a181dacb",
|
"id": "993efa8d",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [],
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"import autograd.numpy as np\n",
|
"import autograd.numpy as np\n",
|
||||||
@@ -1005,8 +1084,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "26a7fe37",
|
"id": "011c734c",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"Below we give a short example of how these cost function may be used\n",
|
"Below we give a short example of how these cost function may be used\n",
|
||||||
"to obtain results if you wish to test them out on your own using\n",
|
"to obtain results if you wish to test them out on your own using\n",
|
||||||
@@ -1016,8 +1097,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 7,
|
"execution_count": 7,
|
||||||
"id": "5894bd31",
|
"id": "9e1b97f5",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [
|
"outputs": [
|
||||||
{
|
{
|
||||||
"name": "stdout",
|
"name": "stdout",
|
||||||
@@ -1045,8 +1129,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "71141fab",
|
"id": "4acd87b2",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"### Activation functions\n",
|
"### Activation functions\n",
|
||||||
"\n",
|
"\n",
|
||||||
@@ -1059,8 +1145,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 8,
|
"execution_count": 8,
|
||||||
"id": "433e0886",
|
"id": "befa86ae",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [],
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"import autograd.numpy as np\n",
|
"import autograd.numpy as np\n",
|
||||||
@@ -1114,8 +1203,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "14c9d538",
|
"id": "c20aa75e",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"Below follows a short demonstration of how to use an activation\n",
|
"Below follows a short demonstration of how to use an activation\n",
|
||||||
"function. The derivative of the activation function will be important\n",
|
"function. The derivative of the activation function will be important\n",
|
||||||
@@ -1127,8 +1218,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 9,
|
"execution_count": 9,
|
||||||
"id": "4c3a44ef",
|
"id": "e209f9e5",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [
|
"outputs": [
|
||||||
{
|
{
|
||||||
"name": "stdout",
|
"name": "stdout",
|
||||||
@@ -1166,8 +1260,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "aa7f604e",
|
"id": "524b409b",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"### The Neural Network\n",
|
"### The Neural Network\n",
|
||||||
"\n",
|
"\n",
|
||||||
@@ -1188,8 +1284,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 10,
|
"execution_count": 10,
|
||||||
"id": "585a227b",
|
"id": "083116d3",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [],
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"import math\n",
|
"import math\n",
|
||||||
@@ -1657,8 +1756,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "7de704af",
|
"id": "e4ba55d0",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"Before we make a model, we will quickly generate a dataset we can use\n",
|
"Before we make a model, we will quickly generate a dataset we can use\n",
|
||||||
"for our linear regression problem as shown below"
|
"for our linear regression problem as shown below"
|
||||||
@@ -1667,8 +1768,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 11,
|
"execution_count": 11,
|
||||||
"id": "212b0a35",
|
"id": "c72a22c6",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [],
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"import autograd.numpy as np\n",
|
"import autograd.numpy as np\n",
|
||||||
@@ -1708,8 +1812,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "1ff4ec7b",
|
"id": "25b63b47",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"Now that we have our dataset ready for the regression, we can create\n",
|
"Now that we have our dataset ready for the regression, we can create\n",
|
||||||
"our regressor. Note that with the seed parameter, we can make sure our\n",
|
"our regressor. Note that with the seed parameter, we can make sure our\n",
|
||||||
@@ -1722,8 +1828,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 12,
|
"execution_count": 12,
|
||||||
"id": "57d8eca3",
|
"id": "b6b4e461",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [],
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"input_nodes = X_train.shape[1]\n",
|
"input_nodes = X_train.shape[1]\n",
|
||||||
@@ -1734,8 +1843,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "f1977bae",
|
"id": "d7860b74",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"We then fit our model with our training data using the scheduler of our choice."
|
"We then fit our model with our training data using the scheduler of our choice."
|
||||||
]
|
]
|
||||||
@@ -1743,8 +1854,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 13,
|
"execution_count": 13,
|
||||||
"id": "31ae2a1e",
|
"id": "00522c73",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [
|
"outputs": [
|
||||||
{
|
{
|
||||||
"name": "stdout",
|
"name": "stdout",
|
||||||
@@ -2573,8 +2687,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "96a64da4",
|
"id": "a57bfb12",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"Due to the progress bar we can see the MSE (train_error) throughout\n",
|
"Due to the progress bar we can see the MSE (train_error) throughout\n",
|
||||||
"the FFNN's training. Note that the fit() function has some optional\n",
|
"the FFNN's training. Note that the fit() function has some optional\n",
|
||||||
@@ -2587,8 +2703,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 14,
|
"execution_count": 14,
|
||||||
"id": "246499fc",
|
"id": "35259e41",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [
|
"outputs": [
|
||||||
{
|
{
|
||||||
"name": "stdout",
|
"name": "stdout",
|
||||||
@@ -10616,8 +10735,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "61d11e17",
|
"id": "4a403370",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"We see that given more epochs to train on, the regressor reaches a lower MSE.\n",
|
"We see that given more epochs to train on, the regressor reaches a lower MSE.\n",
|
||||||
"\n",
|
"\n",
|
||||||
@@ -10629,8 +10750,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 15,
|
"execution_count": 15,
|
||||||
"id": "b5b91899",
|
"id": "6c791130",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [],
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"from sklearn.datasets import load_breast_cancer\n",
|
"from sklearn.datasets import load_breast_cancer\n",
|
||||||
@@ -10652,8 +10776,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 16,
|
"execution_count": 16,
|
||||||
"id": "4bd01b3d",
|
"id": "0ac0258d",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [],
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"input_nodes = X_train.shape[1]\n",
|
"input_nodes = X_train.shape[1]\n",
|
||||||
@@ -10664,8 +10791,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "3a138d66",
|
"id": "a9c74baa",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"We will now make use of our validation data by passing it into our fit function as a keyword argument"
|
"We will now make use of our validation data by passing it into our fit function as a keyword argument"
|
||||||
]
|
]
|
||||||
@@ -10673,8 +10802,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 17,
|
"execution_count": 17,
|
||||||
"id": "faf1534e",
|
"id": "e5021bbf",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [
|
"outputs": [
|
||||||
{
|
{
|
||||||
"name": "stdout",
|
"name": "stdout",
|
||||||
@@ -18703,8 +18835,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "33d2d346",
|
"id": "ae0148bb",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"Finally, we will create a neural network with 2 hidden layers with activation functions."
|
"Finally, we will create a neural network with 2 hidden layers with activation functions."
|
||||||
]
|
]
|
||||||
@@ -18712,8 +18846,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 18,
|
"execution_count": 18,
|
||||||
"id": "a8f0a61c",
|
"id": "bb7e4340",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [],
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"input_nodes = X_train.shape[1]\n",
|
"input_nodes = X_train.shape[1]\n",
|
||||||
@@ -18729,14 +18866,23 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 19,
|
"execution_count": 19,
|
||||||
"id": "d651702d",
|
"id": "d0655afb",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [
|
"outputs": [
|
||||||
{
|
{
|
||||||
"name": "stdout",
|
"name": "stdout",
|
||||||
"output_type": "stream",
|
"output_type": "stream",
|
||||||
"text": [
|
"text": [
|
||||||
"Adam: Eta=0.0001, Lambda=0\n",
|
"Adam: Eta=0.0001, Lambda=0\n"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "stdout",
|
||||||
|
"output_type": "stream",
|
||||||
|
"text": [
|
||||||
"\r",
|
"\r",
|
||||||
" [----------------------------------------] 0.000% | train_error: 11.3 | train_acc: 0.453 | val_error: 10.4 | val_acc: 0.497 "
|
" [----------------------------------------] 0.000% | train_error: 11.3 | train_acc: 0.453 | val_error: 10.4 | val_acc: 0.497 "
|
||||||
]
|
]
|
||||||
@@ -26759,8 +26905,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "200b64c0",
|
"id": "dc8a90af",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"### Multiclass classification\n",
|
"### Multiclass classification\n",
|
||||||
"\n",
|
"\n",
|
||||||
@@ -26772,8 +26920,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 20,
|
"execution_count": 20,
|
||||||
"id": "98b8d59b",
|
"id": "9305c08a",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [
|
"outputs": [
|
||||||
{
|
{
|
||||||
"name": "stdout",
|
"name": "stdout",
|
||||||
@@ -34830,8 +34981,10 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "cadd7e47",
|
"id": "8e208051",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"## Testing the XOR gate and other gates\n",
|
"## Testing the XOR gate and other gates\n",
|
||||||
"\n",
|
"\n",
|
||||||
@@ -34841,8 +34994,11 @@
|
|||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"execution_count": 21,
|
"execution_count": 21,
|
||||||
"id": "a40e5cfe",
|
"id": "db016d41",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"collapsed": false,
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"outputs": [
|
"outputs": [
|
||||||
{
|
{
|
||||||
"name": "stdout",
|
"name": "stdout",
|
||||||
@@ -42879,19 +43035,16 @@
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"id": "b635db73",
|
"id": "d1384449",
|
||||||
"metadata": {},
|
"metadata": {
|
||||||
|
"editable": true
|
||||||
|
},
|
||||||
"source": [
|
"source": [
|
||||||
"Not bad, but the results depend strongly on the learning reate. Try different learning rates."
|
"Not bad, but the results depend strongly on the learning reate. Try different learning rates."
|
||||||
]
|
]
|
||||||
}
|
}
|
||||||
],
|
],
|
||||||
"metadata": {
|
"metadata": {
|
||||||
"kernelspec": {
|
|
||||||
"display_name": "Python 3 (ipykernel)",
|
|
||||||
"language": "python",
|
|
||||||
"name": "python3"
|
|
||||||
},
|
|
||||||
"language_info": {
|
"language_info": {
|
||||||
"codemirror_mode": {
|
"codemirror_mode": {
|
||||||
"name": "ipython",
|
"name": "ipython",
|
||||||
|
|||||||
@@ -160,13 +160,14 @@
|
|||||||
# and multiplying with $\boldsymbol{a}_h$ gives the output
|
# and multiplying with $\boldsymbol{a}_h$ gives the output
|
||||||
|
|
||||||
# $$
|
# $$
|
||||||
# \boldsymbol{a}_o=\boldsymbol{w}_h^T\begin{bmatrix} 0 & 0 \\
|
# \boldsymbol{a}_o=\begin{bmatrix} 0 & 0 \\
|
||||||
# 1 & 0 \\
|
# 1 & 0 \\
|
||||||
# 1 & 0 \\
|
# 1 & 0 \\
|
||||||
# 2 & 1 \end{bmatrix}=\begin{bmatrix} 0 \\ 1 \\ 1 \\0\end{bmatrix},
|
# 2 & 1 \end{bmatrix}\begin{bmatrix} 1 \\
|
||||||
|
# -2\end{bmatrix}=\begin{bmatrix} 0 \\ 1 \\ 1 \\0\end{bmatrix},
|
||||||
# $$
|
# $$
|
||||||
|
|
||||||
# the wanted result.
|
# the wanted result. Pay attention to the dimensionalities as well.
|
||||||
|
|
||||||
# ## Setting up the Neural Network
|
# ## Setting up the Neural Network
|
||||||
#
|
#
|
||||||
|
|||||||
File diff suppressed because it is too large
Load Diff
@@ -210,13 +210,14 @@
|
|||||||
# and multiplying with $\boldsymbol{a}_h$ gives the output
|
# and multiplying with $\boldsymbol{a}_h$ gives the output
|
||||||
|
|
||||||
# $$
|
# $$
|
||||||
# \boldsymbol{a}_o=\boldsymbol{w}_h^T\begin{bmatrix} 0 & 0 \\
|
# \boldsymbol{a}_o=\begin{bmatrix} 0 & 0 \\
|
||||||
# 1 & 0 \\
|
# 1 & 0 \\
|
||||||
# 1 & 0 \\
|
# 1 & 0 \\
|
||||||
# 2 & 1 \end{bmatrix}=\begin{bmatrix} 0 \\ 1 \\ 1 \\0\end{bmatrix},
|
# 2 & 1 \end{bmatrix}\begin{bmatrix} 1 \\
|
||||||
|
# -2\end{bmatrix}=\begin{bmatrix} 0 \\ 1 \\ 1 \\0\end{bmatrix},
|
||||||
# $$
|
# $$
|
||||||
|
|
||||||
# the wanted result.
|
# the wanted result. Pay attention to the dimensionalities as well.
|
||||||
|
|
||||||
# ## Setting up the Neural Network
|
# ## Setting up the Neural Network
|
||||||
#
|
#
|
||||||
|
|||||||
+291
-290
File diff suppressed because it is too large
Load Diff
Reference in New Issue
Block a user