2460 lines
79 KiB
Plaintext
2460 lines
79 KiB
Plaintext
{
|
|
"cells": [
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "2303c986",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"<!-- HTML file automatically generated from DocOnce source (https://github.com/doconce/doconce/)\n",
|
|
"doconce format html week40.do.txt --no_mako -->\n",
|
|
"<!-- dom:TITLE: Week 40: Gradient descent methods (continued) and start Neural networks -->"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "75c3b33e",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"# Week 40: Gradient descent methods (continued) and start Neural networks\n",
|
|
"**Morten Hjorth-Jensen**, Department of Physics, University of Oslo, Norway\n",
|
|
"\n",
|
|
"Date: **September 29-October 3, 2025**"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "4ba50982",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## Lecture Monday September 29, 2025\n",
|
|
"1. Logistic regression and gradient descent, examples on how to code\n",
|
|
"<!-- o Automatic differentiation and gradient descent, examples using Logistic regression -->\n",
|
|
"\n",
|
|
"2. Start with the basics of Neural Networks, setting up the basic steps, from the simple perceptron model to the multi-layer perceptron model\n",
|
|
"\n",
|
|
"3. Video of lecture at <https://youtu.be/MS3Tv8FVArs>\n",
|
|
"\n",
|
|
"4. Whiteboard notes at <https://github.com/CompPhysics/MachineLearning/blob/master/doc/HandWrittenNotes/2025/FYSSTKweek40.pdf>"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "1d527020",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## Suggested readings and videos\n",
|
|
"**Readings and Videos:**\n",
|
|
"\n",
|
|
"1. The lecture notes for week 40 (these notes)\n",
|
|
"<!-- o For a good discussion on gradient methods, we would like to recommend Goodfellow et al section 4.3-4.5 and# sections 8.3-8.6. We will come back to the latter chapter in our discussion of Neural networks as well. -->\n",
|
|
"\n",
|
|
"2. For neural networks we recommend Goodfellow et al chapter 6 and Raschka et al chapter 2 (contains also material about gradient descent) and chapter 11 (we will use this next week)\n",
|
|
"<!-- o Video on gradient descent at <https://www.youtube.com/watch?v=sDv4f4s2SB8> -->\n",
|
|
"<!-- o Video on automatic differentiation at <https://www.youtube.com/watch?v=wG_nF1awSSY> -->\n",
|
|
"\n",
|
|
"3. Neural Networks demystified at <https://www.youtube.com/watch?v=bxe2T-V8XRs&list=PLiaHhY2iBX9hdHaRr6b7XevZtgZRa1PoU&ab_channel=WelchLabs>\n",
|
|
"\n",
|
|
"4. Building Neural Networks from scratch at URL:https://www.youtube.com/watch?v=Wo5dMEP_BbI&list=PLQVvvaa0QuDcjD5BAw2DxE6OF2tius3V3&ab_channel=sentdex\""
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "63a4d497",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## Lab sessions Tuesday and Wednesday\n",
|
|
"**Material for the active learning sessions on Tuesday and Wednesday.**\n",
|
|
"\n",
|
|
" * Work on project 1 and discussions on how to structure your report\n",
|
|
"\n",
|
|
" * No weekly exercises for week 40, project work only\n",
|
|
"\n",
|
|
" * Video on how to write scientific reports recorded during one of the lab sessions at <https://youtu.be/tVW1ZDmZnwM>\n",
|
|
"\n",
|
|
" * A general guideline can be found at <https://github.com/CompPhysics/MachineLearning/blob/master/doc/Projects/EvaluationGrading/EvaluationForm.md>."
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "73621d6b",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## Logistic Regression, from last week\n",
|
|
"\n",
|
|
"In linear regression our main interest was centered on learning the\n",
|
|
"coefficients of a functional fit (say a polynomial) in order to be\n",
|
|
"able to predict the response of a continuous variable on some unseen\n",
|
|
"data. The fit to the continuous variable $y_i$ is based on some\n",
|
|
"independent variables $\\boldsymbol{x}_i$. Linear regression resulted in\n",
|
|
"analytical expressions for standard ordinary Least Squares or Ridge\n",
|
|
"regression (in terms of matrices to invert) for several quantities,\n",
|
|
"ranging from the variance and thereby the confidence intervals of the\n",
|
|
"parameters $\\boldsymbol{\\theta}$ to the mean squared error. If we can invert\n",
|
|
"the product of the design matrices, linear regression gives then a\n",
|
|
"simple recipe for fitting our data."
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "fc1df17b",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## Classification problems\n",
|
|
"\n",
|
|
"Classification problems, however, are concerned with outcomes taking\n",
|
|
"the form of discrete variables (i.e. categories). We may for example,\n",
|
|
"on the basis of DNA sequencing for a number of patients, like to find\n",
|
|
"out which mutations are important for a certain disease; or based on\n",
|
|
"scans of various patients' brains, figure out if there is a tumor or\n",
|
|
"not; or given a specific physical system, we'd like to identify its\n",
|
|
"state, say whether it is an ordered or disordered system (typical\n",
|
|
"situation in solid state physics); or classify the status of a\n",
|
|
"patient, whether she/he has a stroke or not and many other similar\n",
|
|
"situations.\n",
|
|
"\n",
|
|
"The most common situation we encounter when we apply logistic\n",
|
|
"regression is that of two possible outcomes, normally denoted as a\n",
|
|
"binary outcome, true or false, positive or negative, success or\n",
|
|
"failure etc."
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "a3d311e6",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## Optimization and Deep learning\n",
|
|
"\n",
|
|
"Logistic regression will also serve as our stepping stone towards\n",
|
|
"neural network algorithms and supervised deep learning. For logistic\n",
|
|
"learning, the minimization of the cost function leads to a non-linear\n",
|
|
"equation in the parameters $\\boldsymbol{\\theta}$. The optimization of the\n",
|
|
"problem calls therefore for minimization algorithms.\n",
|
|
"\n",
|
|
"As we have discussed earlier, this forms the\n",
|
|
"bottle neck of all machine learning algorithms, namely how to find\n",
|
|
"reliable minima of a multi-variable function. This leads us to the\n",
|
|
"family of gradient descent methods. The latter are the working horses\n",
|
|
"of basically all modern machine learning algorithms.\n",
|
|
"\n",
|
|
"We note also that many of the topics discussed here on logistic \n",
|
|
"regression are also commonly used in modern supervised Deep Learning\n",
|
|
"models, as we will see later."
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "4120d6f9",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## Basics\n",
|
|
"\n",
|
|
"We consider the case where the outputs/targets, also called the\n",
|
|
"responses or the outcomes, $y_i$ are discrete and only take values\n",
|
|
"from $k=0,\\dots,K-1$ (i.e. $K$ classes).\n",
|
|
"\n",
|
|
"The goal is to predict the\n",
|
|
"output classes from the design matrix $\\boldsymbol{X}\\in\\mathbb{R}^{n\\times p}$\n",
|
|
"made of $n$ samples, each of which carries $p$ features or predictors. The\n",
|
|
"primary goal is to identify the classes to which new unseen samples\n",
|
|
"belong.\n",
|
|
"\n",
|
|
"Last week we specialized to the case of two classes only, with outputs\n",
|
|
"$y_i=0$ and $y_i=1$. Our outcomes could represent the status of a\n",
|
|
"credit card user that could default or not on her/his credit card\n",
|
|
"debt. That is"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "9e85d1e4",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"$$\n",
|
|
"y_i = \\begin{bmatrix} 0 & \\mathrm{no}\\\\ 1 & \\mathrm{yes} \\end{bmatrix}.\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "a0d8c838",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## Two parameters\n",
|
|
"\n",
|
|
"We assume now that we have two classes with $y_i$ either $0$ or $1$. Furthermore we assume also that we have only two parameters $\\theta$ in our fitting of the Sigmoid function, that is we define probabilities"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "7cea7945",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"$$\n",
|
|
"\\begin{align*}\n",
|
|
"p(y_i=1|x_i,\\boldsymbol{\\theta}) &= \\frac{\\exp{(\\theta_0+\\theta_1x_i)}}{1+\\exp{(\\theta_0+\\theta_1x_i)}},\\nonumber\\\\\n",
|
|
"p(y_i=0|x_i,\\boldsymbol{\\theta}) &= 1 - p(y_i=1|x_i,\\boldsymbol{\\theta}),\n",
|
|
"\\end{align*}\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "6adc5106",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"where $\\boldsymbol{\\theta}$ are the weights we wish to extract from data, in our case $\\theta_0$ and $\\theta_1$. \n",
|
|
"\n",
|
|
"Note that we used"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "f976068e",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"$$\n",
|
|
"p(y_i=0\\vert x_i, \\boldsymbol{\\theta}) = 1-p(y_i=1\\vert x_i, \\boldsymbol{\\theta}).\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "dedf9f0e",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## Maximum likelihood\n",
|
|
"\n",
|
|
"In order to define the total likelihood for all possible outcomes from a \n",
|
|
"dataset $\\mathcal{D}=\\{(y_i,x_i)\\}$, with the binary labels\n",
|
|
"$y_i\\in\\{0,1\\}$ and where the data points are drawn independently, we use the so-called [Maximum Likelihood Estimation](https://en.wikipedia.org/wiki/Maximum_likelihood_estimation) (MLE) principle. \n",
|
|
"We aim thus at maximizing \n",
|
|
"the probability of seeing the observed data. We can then approximate the \n",
|
|
"likelihood in terms of the product of the individual probabilities of a specific outcome $y_i$, that is"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "bd8b54ab",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"$$\n",
|
|
"\\begin{align*}\n",
|
|
"P(\\mathcal{D}|\\boldsymbol{\\theta})& = \\prod_{i=1}^n \\left[p(y_i=1|x_i,\\boldsymbol{\\theta})\\right]^{y_i}\\left[1-p(y_i=1|x_i,\\boldsymbol{\\theta}))\\right]^{1-y_i}\\nonumber \\\\\n",
|
|
"\\end{align*}\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "57bfb17f",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"from which we obtain the log-likelihood and our **cost/loss** function"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "00aee268",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"$$\n",
|
|
"\\mathcal{C}(\\boldsymbol{\\theta}) = \\sum_{i=1}^n \\left( y_i\\log{p(y_i=1|x_i,\\boldsymbol{\\theta})} + (1-y_i)\\log\\left[1-p(y_i=1|x_i,\\boldsymbol{\\theta}))\\right]\\right).\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "e12940f3",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## The cost function rewritten\n",
|
|
"\n",
|
|
"Reordering the logarithms, we can rewrite the **cost/loss** function as"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "e5b2b29e",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"$$\n",
|
|
"\\mathcal{C}(\\boldsymbol{\\theta}) = \\sum_{i=1}^n \\left(y_i(\\theta_0+\\theta_1x_i) -\\log{(1+\\exp{(\\theta_0+\\theta_1x_i)})}\\right).\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "c6c0ba4c",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"The maximum likelihood estimator is defined as the set of parameters that maximize the log-likelihood where we maximize with respect to $\\theta$.\n",
|
|
"Since the cost (error) function is just the negative log-likelihood, for logistic regression we have that"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "46ee2ea8",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"$$\n",
|
|
"\\mathcal{C}(\\boldsymbol{\\theta})=-\\sum_{i=1}^n \\left(y_i(\\theta_0+\\theta_1x_i) -\\log{(1+\\exp{(\\theta_0+\\theta_1x_i)})}\\right).\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "9a05709b",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"This equation is known in statistics as the **cross entropy**. Finally, we note that just as in linear regression, \n",
|
|
"in practice we often supplement the cross-entropy with additional regularization terms, usually $L_1$ and $L_2$ regularization as we did for Ridge and Lasso regression."
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "ae1362c9",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## Minimizing the cross entropy\n",
|
|
"\n",
|
|
"The cross entropy is a convex function of the weights $\\boldsymbol{\\theta}$ and,\n",
|
|
"therefore, any local minimizer is a global minimizer. \n",
|
|
"\n",
|
|
"Minimizing this\n",
|
|
"cost function with respect to the two parameters $\\theta_0$ and $\\theta_1$ we obtain"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "57f4670b",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"$$\n",
|
|
"\\frac{\\partial \\mathcal{C}(\\boldsymbol{\\theta})}{\\partial \\theta_0} = -\\sum_{i=1}^n \\left(y_i -\\frac{\\exp{(\\theta_0+\\theta_1x_i)}}{1+\\exp{(\\theta_0+\\theta_1x_i)}}\\right),\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "1dc19f59",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"and"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "4e96dc87",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"$$\n",
|
|
"\\frac{\\partial \\mathcal{C}(\\boldsymbol{\\theta})}{\\partial \\theta_1} = -\\sum_{i=1}^n \\left(y_ix_i -x_i\\frac{\\exp{(\\theta_0+\\theta_1x_i)}}{1+\\exp{(\\theta_0+\\theta_1x_i)}}\\right).\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "fa77bec9",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## A more compact expression\n",
|
|
"\n",
|
|
"Let us now define a vector $\\boldsymbol{y}$ with $n$ elements $y_i$, an\n",
|
|
"$n\\times p$ matrix $\\boldsymbol{X}$ which contains the $x_i$ values and a\n",
|
|
"vector $\\boldsymbol{p}$ of fitted probabilities $p(y_i\\vert x_i,\\boldsymbol{\\theta})$. We can rewrite in a more compact form the first\n",
|
|
"derivative of the cost function as"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "1b013fd2",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"$$\n",
|
|
"\\frac{\\partial \\mathcal{C}(\\boldsymbol{\\theta})}{\\partial \\boldsymbol{\\theta}} = -\\boldsymbol{X}^T\\left(\\boldsymbol{y}-\\boldsymbol{p}\\right).\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "910f36dd",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"If we in addition define a diagonal matrix $\\boldsymbol{W}$ with elements \n",
|
|
"$p(y_i\\vert x_i,\\boldsymbol{\\theta})(1-p(y_i\\vert x_i,\\boldsymbol{\\theta})$, we can obtain a compact expression of the second derivative as"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "8212d0ed",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"$$\n",
|
|
"\\frac{\\partial^2 \\mathcal{C}(\\boldsymbol{\\theta})}{\\partial \\boldsymbol{\\theta}\\partial \\boldsymbol{\\theta}^T} = \\boldsymbol{X}^T\\boldsymbol{W}\\boldsymbol{X}.\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "7ae7078b",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## Extending to more predictors\n",
|
|
"\n",
|
|
"Within a binary classification problem, we can easily expand our model to include multiple predictors. Our ratio between likelihoods is then with $p$ predictors"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "59e57d7c",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"$$\n",
|
|
"\\log{ \\frac{p(\\boldsymbol{\\theta}\\boldsymbol{x})}{1-p(\\boldsymbol{\\theta}\\boldsymbol{x})}} = \\theta_0+\\theta_1x_1+\\theta_2x_2+\\dots+\\theta_px_p.\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "6ffe0955",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"Here we defined $\\boldsymbol{x}=[1,x_1,x_2,\\dots,x_p]$ and $\\boldsymbol{\\theta}=[\\theta_0, \\theta_1, \\dots, \\theta_p]$ leading to"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "56e9bd82",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"$$\n",
|
|
"p(\\boldsymbol{\\theta}\\boldsymbol{x})=\\frac{ \\exp{(\\theta_0+\\theta_1x_1+\\theta_2x_2+\\dots+\\theta_px_p)}}{1+\\exp{(\\theta_0+\\theta_1x_1+\\theta_2x_2+\\dots+\\theta_px_p)}}.\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "86b12946",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## Including more classes\n",
|
|
"\n",
|
|
"Till now we have mainly focused on two classes, the so-called binary\n",
|
|
"system. Suppose we wish to extend to $K$ classes. Let us for the sake\n",
|
|
"of simplicity assume we have only two predictors. We have then following model"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "d55394df",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"$$\n",
|
|
"\\log{\\frac{p(C=1\\vert x)}{p(K\\vert x)}} = \\theta_{10}+\\theta_{11}x_1,\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "ee01378a",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"and"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "c7fadfbb",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"$$\n",
|
|
"\\log{\\frac{p(C=2\\vert x)}{p(K\\vert x)}} = \\theta_{20}+\\theta_{21}x_1,\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "e8310f63",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"and so on till the class $C=K-1$ class"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "be651647",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"$$\n",
|
|
"\\log{\\frac{p(C=K-1\\vert x)}{p(K\\vert x)}} = \\theta_{(K-1)0}+\\theta_{(K-1)1}x_1,\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "e277c601",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"and the model is specified in term of $K-1$ so-called log-odds or\n",
|
|
"**logit** transformations."
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "aea3a410",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## More classes\n",
|
|
"\n",
|
|
"In our discussion of neural networks we will encounter the above again\n",
|
|
"in terms of a slightly modified function, the so-called **Softmax** function.\n",
|
|
"\n",
|
|
"The softmax function is used in various multiclass classification\n",
|
|
"methods, such as multinomial logistic regression (also known as\n",
|
|
"softmax regression), multiclass linear discriminant analysis, naive\n",
|
|
"Bayes classifiers, and artificial neural networks. Specifically, in\n",
|
|
"multinomial logistic regression and linear discriminant analysis, the\n",
|
|
"input to the function is the result of $K$ distinct linear functions,\n",
|
|
"and the predicted probability for the $k$-th class given a sample\n",
|
|
"vector $\\boldsymbol{x}$ and a weighting vector $\\boldsymbol{\\theta}$ is (with two\n",
|
|
"predictors):"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "bfa7221f",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"$$\n",
|
|
"p(C=k\\vert \\mathbf {x} )=\\frac{\\exp{(\\theta_{k0}+\\theta_{k1}x_1)}}{1+\\sum_{l=1}^{K-1}\\exp{(\\theta_{l0}+\\theta_{l1}x_1)}}.\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "3d749c39",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"It is easy to extend to more predictors. The final class is"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "dc061a39",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"$$\n",
|
|
"p(C=K\\vert \\mathbf {x} )=\\frac{1}{1+\\sum_{l=1}^{K-1}\\exp{(\\theta_{l0}+\\theta_{l1}x_1)}},\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "8ea10488",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"and they sum to one. Our earlier discussions were all specialized to\n",
|
|
"the case with two classes only. It is easy to see from the above that\n",
|
|
"what we derived earlier is compatible with these equations.\n",
|
|
"\n",
|
|
"To find the optimal parameters we would typically use a gradient\n",
|
|
"descent method. Newton's method and gradient descent methods are\n",
|
|
"discussed in the material on [optimization\n",
|
|
"methods](https://compphysics.github.io/MachineLearning/doc/pub/Splines/html/Splines-bs.html)."
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "9cb3baf8",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## Optimization, the central part of any Machine Learning algortithm\n",
|
|
"\n",
|
|
"Almost every problem in machine learning and data science starts with\n",
|
|
"a dataset $X$, a model $g(\\theta)$, which is a function of the\n",
|
|
"parameters $\\theta$ and a cost function $C(X, g(\\theta))$ that allows\n",
|
|
"us to judge how well the model $g(\\theta)$ explains the observations\n",
|
|
"$X$. The model is fit by finding the values of $\\theta$ that minimize\n",
|
|
"the cost function. Ideally we would be able to solve for $\\theta$\n",
|
|
"analytically, however this is not possible in general and we must use\n",
|
|
"some approximative/numerical method to compute the minimum."
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "387393d7",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## Revisiting our Logistic Regression case\n",
|
|
"\n",
|
|
"In our discussion on Logistic Regression we studied the \n",
|
|
"case of\n",
|
|
"two classes, with $y_i$ either\n",
|
|
"$0$ or $1$. Furthermore we assumed also that we have only two\n",
|
|
"parameters $\\theta$ in our fitting, that is we\n",
|
|
"defined probabilities"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "30f64659",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"$$\n",
|
|
"\\begin{align*}\n",
|
|
"p(y_i=1|x_i,\\boldsymbol{\\theta}) &= \\frac{\\exp{(\\theta_0+\\theta_1x_i)}}{1+\\exp{(\\theta_0+\\theta_1x_i)}},\\nonumber\\\\\n",
|
|
"p(y_i=0|x_i,\\boldsymbol{\\theta}) &= 1 - p(y_i=1|x_i,\\boldsymbol{\\theta}),\n",
|
|
"\\end{align*}\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "3ba65422",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"where $\\boldsymbol{\\theta}$ are the weights we wish to extract from data, in our case $\\theta_0$ and $\\theta_1$."
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "005f46d7",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## The equations to solve\n",
|
|
"\n",
|
|
"Our compact equations used a definition of a vector $\\boldsymbol{y}$ with $n$\n",
|
|
"elements $y_i$, an $n\\times p$ matrix $\\boldsymbol{X}$ which contains the\n",
|
|
"$x_i$ values and a vector $\\boldsymbol{p}$ of fitted probabilities\n",
|
|
"$p(y_i\\vert x_i,\\boldsymbol{\\theta})$. We rewrote in a more compact form\n",
|
|
"the first derivative of the cost function as"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "61a638bc",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"$$\n",
|
|
"\\frac{\\partial \\mathcal{C}(\\boldsymbol{\\theta})}{\\partial \\boldsymbol{\\theta}} = -\\boldsymbol{X}^T\\left(\\boldsymbol{y}-\\boldsymbol{p}\\right).\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "469c0042",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"If we in addition define a diagonal matrix $\\boldsymbol{W}$ with elements \n",
|
|
"$p(y_i\\vert x_i,\\boldsymbol{\\theta})(1-p(y_i\\vert x_i,\\boldsymbol{\\theta})$, we can obtain a compact expression of the second derivative as"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "0af5449a",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"$$\n",
|
|
"\\frac{\\partial^2 \\mathcal{C}(\\boldsymbol{\\theta})}{\\partial \\boldsymbol{\\theta}\\partial \\boldsymbol{\\theta}^T} = \\boldsymbol{X}^T\\boldsymbol{W}\\boldsymbol{X}.\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "f4c16b4f",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"This defines what is called the Hessian matrix."
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "ddbe7f50",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## Solving using Newton-Raphson's method\n",
|
|
"\n",
|
|
"If we can set up these equations, Newton-Raphson's iterative method is normally the method of choice. It requires however that we can compute in an efficient way the matrices that define the first and second derivatives. \n",
|
|
"\n",
|
|
"Our iterative scheme is then given by"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "52830f96",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"$$\n",
|
|
"\\boldsymbol{\\theta}^{\\mathrm{new}} = \\boldsymbol{\\theta}^{\\mathrm{old}}-\\left(\\frac{\\partial^2 \\mathcal{C}(\\boldsymbol{\\theta})}{\\partial \\boldsymbol{\\theta}\\partial \\boldsymbol{\\theta}^T}\\right)^{-1}_{\\boldsymbol{\\theta}^{\\mathrm{old}}}\\times \\left(\\frac{\\partial \\mathcal{C}(\\boldsymbol{\\theta})}{\\partial \\boldsymbol{\\theta}}\\right)_{\\boldsymbol{\\theta}^{\\mathrm{old}}},\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "1b8a1c14",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"or in matrix form as"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "8ad73cea",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"$$\n",
|
|
"\\boldsymbol{\\theta}^{\\mathrm{new}} = \\boldsymbol{\\theta}^{\\mathrm{old}}-\\left(\\boldsymbol{X}^T\\boldsymbol{W}\\boldsymbol{X} \\right)^{-1}\\times \\left(-\\boldsymbol{X}^T(\\boldsymbol{y}-\\boldsymbol{p}) \\right)_{\\boldsymbol{\\theta}^{\\mathrm{old}}}.\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "6d47dd0b",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"The right-hand side is computed with the old values of $\\theta$. \n",
|
|
"\n",
|
|
"If we can compute these matrices, in particular the Hessian, the above is often the easiest method to implement."
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "f399c2f4",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## Example code for Logistic Regression\n",
|
|
"\n",
|
|
"Here we make a class for Logistic regression. The code uses a simple data set and includes both a binary case and a multiclass case."
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "code",
|
|
"execution_count": 1,
|
|
"id": "79f6b6fc",
|
|
"metadata": {
|
|
"collapsed": false,
|
|
"editable": true
|
|
},
|
|
"outputs": [],
|
|
"source": [
|
|
"import numpy as np\n",
|
|
"\n",
|
|
"class LogisticRegression:\n",
|
|
" \"\"\"\n",
|
|
" Logistic Regression for binary and multiclass classification.\n",
|
|
" \"\"\"\n",
|
|
" def __init__(self, lr=0.01, epochs=1000, fit_intercept=True, verbose=False):\n",
|
|
" self.lr = lr # Learning rate for gradient descent\n",
|
|
" self.epochs = epochs # Number of iterations\n",
|
|
" self.fit_intercept = fit_intercept # Whether to add intercept (bias)\n",
|
|
" self.verbose = verbose # Print loss during training if True\n",
|
|
" self.weights = None\n",
|
|
" self.multi_class = False # Will be determined at fit time\n",
|
|
"\n",
|
|
" def _add_intercept(self, X):\n",
|
|
" \"\"\"Add intercept term (column of ones) to feature matrix.\"\"\"\n",
|
|
" intercept = np.ones((X.shape[0], 1))\n",
|
|
" return np.concatenate((intercept, X), axis=1)\n",
|
|
"\n",
|
|
" def _sigmoid(self, z):\n",
|
|
" \"\"\"Sigmoid function for binary logistic.\"\"\"\n",
|
|
" return 1 / (1 + np.exp(-z))\n",
|
|
"\n",
|
|
" def _softmax(self, Z):\n",
|
|
" \"\"\"Softmax function for multiclass logistic.\"\"\"\n",
|
|
" exp_Z = np.exp(Z - np.max(Z, axis=1, keepdims=True))\n",
|
|
" return exp_Z / np.sum(exp_Z, axis=1, keepdims=True)\n",
|
|
"\n",
|
|
" def fit(self, X, y):\n",
|
|
" \"\"\"\n",
|
|
" Train the logistic regression model using gradient descent.\n",
|
|
" Supports binary (sigmoid) and multiclass (softmax) based on y.\n",
|
|
" \"\"\"\n",
|
|
" X = np.array(X)\n",
|
|
" y = np.array(y)\n",
|
|
" n_samples, n_features = X.shape\n",
|
|
"\n",
|
|
" # Add intercept if needed\n",
|
|
" if self.fit_intercept:\n",
|
|
" X = self._add_intercept(X)\n",
|
|
" n_features += 1\n",
|
|
"\n",
|
|
" # Determine classes and mode (binary vs multiclass)\n",
|
|
" unique_classes = np.unique(y)\n",
|
|
" if len(unique_classes) > 2:\n",
|
|
" self.multi_class = True\n",
|
|
" else:\n",
|
|
" self.multi_class = False\n",
|
|
"\n",
|
|
" # ----- Multiclass case -----\n",
|
|
" if self.multi_class:\n",
|
|
" n_classes = len(unique_classes)\n",
|
|
" # Map original labels to 0...n_classes-1\n",
|
|
" class_to_index = {c: idx for idx, c in enumerate(unique_classes)}\n",
|
|
" y_indices = np.array([class_to_index[c] for c in y])\n",
|
|
" # Initialize weight matrix (features x classes)\n",
|
|
" self.weights = np.zeros((n_features, n_classes))\n",
|
|
"\n",
|
|
" # One-hot encode y\n",
|
|
" Y_onehot = np.zeros((n_samples, n_classes))\n",
|
|
" Y_onehot[np.arange(n_samples), y_indices] = 1\n",
|
|
"\n",
|
|
" # Gradient descent\n",
|
|
" for epoch in range(self.epochs):\n",
|
|
" scores = X.dot(self.weights) # Linear scores (n_samples x n_classes)\n",
|
|
" probs = self._softmax(scores) # Probabilities (n_samples x n_classes)\n",
|
|
" # Compute gradient (features x classes)\n",
|
|
" gradient = (1 / n_samples) * X.T.dot(probs - Y_onehot)\n",
|
|
" # Update weights\n",
|
|
" self.weights -= self.lr * gradient\n",
|
|
"\n",
|
|
" if self.verbose and epoch % 100 == 0:\n",
|
|
" # Compute current loss (categorical cross-entropy)\n",
|
|
" loss = -np.sum(Y_onehot * np.log(probs + 1e-15)) / n_samples\n",
|
|
" print(f\"[Epoch {epoch}] Multiclass loss: {loss:.4f}\")\n",
|
|
"\n",
|
|
" # ----- Binary case -----\n",
|
|
" else:\n",
|
|
" # Convert y to 0/1 if not already\n",
|
|
" if not np.array_equal(unique_classes, [0, 1]):\n",
|
|
" # Map the two classes to 0 and 1\n",
|
|
" class0, class1 = unique_classes\n",
|
|
" y_binary = np.where(y == class1, 1, 0)\n",
|
|
" else:\n",
|
|
" y_binary = y.copy().astype(int)\n",
|
|
"\n",
|
|
" # Initialize weights vector (features,)\n",
|
|
" self.weights = np.zeros(n_features)\n",
|
|
"\n",
|
|
" # Gradient descent\n",
|
|
" for epoch in range(self.epochs):\n",
|
|
" linear_model = X.dot(self.weights) # (n_samples,)\n",
|
|
" probs = self._sigmoid(linear_model) # (n_samples,)\n",
|
|
" # Gradient for binary cross-entropy\n",
|
|
" gradient = (1 / n_samples) * X.T.dot(probs - y_binary)\n",
|
|
" self.weights -= self.lr * gradient\n",
|
|
"\n",
|
|
" if self.verbose and epoch % 100 == 0:\n",
|
|
" # Compute binary cross-entropy loss\n",
|
|
" loss = -np.mean(\n",
|
|
" y_binary * np.log(probs + 1e-15) + \n",
|
|
" (1 - y_binary) * np.log(1 - probs + 1e-15)\n",
|
|
" )\n",
|
|
" print(f\"[Epoch {epoch}] Binary loss: {loss:.4f}\")\n",
|
|
"\n",
|
|
" def predict_prob(self, X):\n",
|
|
" \"\"\"\n",
|
|
" Compute probability estimates. Returns a 1D array for binary or\n",
|
|
" a 2D array (n_samples x n_classes) for multiclass.\n",
|
|
" \"\"\"\n",
|
|
" X = np.array(X)\n",
|
|
" # Add intercept if the model used it\n",
|
|
" if self.fit_intercept:\n",
|
|
" X = self._add_intercept(X)\n",
|
|
" scores = X.dot(self.weights)\n",
|
|
" if self.multi_class:\n",
|
|
" return self._softmax(scores)\n",
|
|
" else:\n",
|
|
" return self._sigmoid(scores)\n",
|
|
"\n",
|
|
" def predict(self, X):\n",
|
|
" \"\"\"\n",
|
|
" Predict class labels for samples in X.\n",
|
|
" Returns integer class labels (0,1 for binary, or 0...C-1 for multiclass).\n",
|
|
" \"\"\"\n",
|
|
" probs = self.predict_prob(X)\n",
|
|
" if self.multi_class:\n",
|
|
" # Choose class with highest probability\n",
|
|
" return np.argmax(probs, axis=1)\n",
|
|
" else:\n",
|
|
" # Threshold at 0.5 for binary\n",
|
|
" return (probs >= 0.5).astype(int)"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "24e84b29",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"The class implements the sigmoid and softmax internally. During fit(),\n",
|
|
"we check the number of classes: if more than 2, we set\n",
|
|
"self.multi_class=True and perform multinomial logistic regression. We\n",
|
|
"one-hot encode the target vector and update a weight matrix with\n",
|
|
"softmax probabilities. Otherwise, we do standard binary logistic\n",
|
|
"regression, converting labels to 0/1 if needed and updating a weight\n",
|
|
"vector. In both cases we use batch gradient descent on the\n",
|
|
"cross-entropy loss (we add a small epsilon 1e-15 to logs for numerical\n",
|
|
"stability). Progress (loss) can be printed if verbose=True."
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "code",
|
|
"execution_count": 2,
|
|
"id": "7a73eca4",
|
|
"metadata": {
|
|
"collapsed": false,
|
|
"editable": true
|
|
},
|
|
"outputs": [],
|
|
"source": [
|
|
"# Evaluation Metrics\n",
|
|
"#We define helper functions for accuracy and cross-entropy loss. Accuracy is the fraction of correct predictions . For loss, we compute the appropriate cross-entropy:\n",
|
|
"\n",
|
|
"def accuracy_score(y_true, y_pred):\n",
|
|
" \"\"\"Accuracy = (# correct predictions) / (total samples).\"\"\"\n",
|
|
" y_true = np.array(y_true)\n",
|
|
" y_pred = np.array(y_pred)\n",
|
|
" return np.mean(y_true == y_pred)\n",
|
|
"\n",
|
|
"def binary_cross_entropy(y_true, y_prob):\n",
|
|
" \"\"\"\n",
|
|
" Binary cross-entropy loss.\n",
|
|
" y_true: true binary labels (0 or 1), y_prob: predicted probabilities for class 1.\n",
|
|
" \"\"\"\n",
|
|
" y_true = np.array(y_true)\n",
|
|
" y_prob = np.clip(np.array(y_prob), 1e-15, 1-1e-15)\n",
|
|
" return -np.mean(y_true * np.log(y_prob) + (1 - y_true) * np.log(1 - y_prob))\n",
|
|
"\n",
|
|
"def categorical_cross_entropy(y_true, y_prob):\n",
|
|
" \"\"\"\n",
|
|
" Categorical cross-entropy loss for multiclass.\n",
|
|
" y_true: true labels (0...C-1), y_prob: array of predicted probabilities (n_samples x C).\n",
|
|
" \"\"\"\n",
|
|
" y_true = np.array(y_true, dtype=int)\n",
|
|
" y_prob = np.clip(np.array(y_prob), 1e-15, 1-1e-15)\n",
|
|
" # One-hot encode true labels\n",
|
|
" n_samples, n_classes = y_prob.shape\n",
|
|
" one_hot = np.zeros_like(y_prob)\n",
|
|
" one_hot[np.arange(n_samples), y_true] = 1\n",
|
|
" # Compute cross-entropy\n",
|
|
" loss_vec = -np.sum(one_hot * np.log(y_prob), axis=1)\n",
|
|
" return np.mean(loss_vec)"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "40d4b30f",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"### Synthetic data generation\n",
|
|
"\n",
|
|
"Binary classification data: Create two Gaussian clusters in 2D. For example, class 0 around mean [-2,-2] and class 1 around [2,2].\n",
|
|
"Multiclass data: Create several Gaussian clusters (one per class) spread out in feature space."
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "code",
|
|
"execution_count": 3,
|
|
"id": "ac0089bf",
|
|
"metadata": {
|
|
"collapsed": false,
|
|
"editable": true
|
|
},
|
|
"outputs": [],
|
|
"source": [
|
|
"import numpy as np\n",
|
|
"\n",
|
|
"def generate_binary_data(n_samples=100, n_features=2, random_state=None):\n",
|
|
" \"\"\"\n",
|
|
" Generate synthetic binary classification data.\n",
|
|
" Returns (X, y) where X is (n_samples x n_features), y in {0,1}.\n",
|
|
" \"\"\"\n",
|
|
" rng = np.random.RandomState(random_state)\n",
|
|
" # Half samples for class 0, half for class 1\n",
|
|
" n0 = n_samples // 2\n",
|
|
" n1 = n_samples - n0\n",
|
|
" # Class 0 around mean -2, class 1 around +2\n",
|
|
" mean0 = -2 * np.ones(n_features)\n",
|
|
" mean1 = 2 * np.ones(n_features)\n",
|
|
" X0 = rng.randn(n0, n_features) + mean0\n",
|
|
" X1 = rng.randn(n1, n_features) + mean1\n",
|
|
" X = np.vstack((X0, X1))\n",
|
|
" y = np.array([0]*n0 + [1]*n1)\n",
|
|
" return X, y\n",
|
|
"\n",
|
|
"def generate_multiclass_data(n_samples=150, n_features=2, n_classes=3, random_state=None):\n",
|
|
" \"\"\"\n",
|
|
" Generate synthetic multiclass data with n_classes Gaussian clusters.\n",
|
|
" \"\"\"\n",
|
|
" rng = np.random.RandomState(random_state)\n",
|
|
" X = []\n",
|
|
" y = []\n",
|
|
" samples_per_class = n_samples // n_classes\n",
|
|
" for cls in range(n_classes):\n",
|
|
" # Random cluster center for each class\n",
|
|
" center = rng.uniform(-5, 5, size=n_features)\n",
|
|
" Xi = rng.randn(samples_per_class, n_features) + center\n",
|
|
" yi = [cls] * samples_per_class\n",
|
|
" X.append(Xi)\n",
|
|
" y.extend(yi)\n",
|
|
" X = np.vstack(X)\n",
|
|
" y = np.array(y)\n",
|
|
" return X, y\n",
|
|
"\n",
|
|
"\n",
|
|
"# Generate and test on binary data\n",
|
|
"X_bin, y_bin = generate_binary_data(n_samples=200, n_features=2, random_state=42)\n",
|
|
"model_bin = LogisticRegression(lr=0.1, epochs=1000)\n",
|
|
"model_bin.fit(X_bin, y_bin)\n",
|
|
"y_prob_bin = model_bin.predict_prob(X_bin) # probabilities for class 1\n",
|
|
"y_pred_bin = model_bin.predict(X_bin) # predicted classes 0 or 1\n",
|
|
"\n",
|
|
"acc_bin = accuracy_score(y_bin, y_pred_bin)\n",
|
|
"loss_bin = binary_cross_entropy(y_bin, y_prob_bin)\n",
|
|
"print(f\"Binary Classification - Accuracy: {acc_bin:.2f}, Cross-Entropy Loss: {loss_bin:.2f}\")\n",
|
|
"#For multiclass:\n",
|
|
"# Generate and test on multiclass data\n",
|
|
"X_multi, y_multi = generate_multiclass_data(n_samples=300, n_features=2, n_classes=3, random_state=1)\n",
|
|
"model_multi = LogisticRegression(lr=0.1, epochs=1000)\n",
|
|
"model_multi.fit(X_multi, y_multi)\n",
|
|
"y_prob_multi = model_multi.predict_prob(X_multi) # (n_samples x 3) probabilities\n",
|
|
"y_pred_multi = model_multi.predict(X_multi) # predicted labels 0,1,2\n",
|
|
"\n",
|
|
"acc_multi = accuracy_score(y_multi, y_pred_multi)\n",
|
|
"loss_multi = categorical_cross_entropy(y_multi, y_prob_multi)\n",
|
|
"print(f\"Multiclass Classification - Accuracy: {acc_multi:.2f}, Cross-Entropy Loss: {loss_multi:.2f}\")\n",
|
|
"\n",
|
|
"# CSV Export\n",
|
|
"import csv\n",
|
|
"\n",
|
|
"# Export binary results\n",
|
|
"with open('binary_results.csv', mode='w', newline='') as f:\n",
|
|
" writer = csv.writer(f)\n",
|
|
" writer.writerow([\"TrueLabel\", \"PredictedLabel\"])\n",
|
|
" for true, pred in zip(y_bin, y_pred_bin):\n",
|
|
" writer.writerow([true, pred])\n",
|
|
"\n",
|
|
"# Export multiclass results\n",
|
|
"with open('multiclass_results.csv', mode='w', newline='') as f:\n",
|
|
" writer = csv.writer(f)\n",
|
|
" writer.writerow([\"TrueLabel\", \"PredictedLabel\"])\n",
|
|
" for true, pred in zip(y_multi, y_pred_multi):\n",
|
|
" writer.writerow([true, pred])"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "1e9acef3",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## Using **Scikit-learn**\n",
|
|
"\n",
|
|
"We show here how we can use a logistic regression case on a data set\n",
|
|
"included in _scikit_learn_, the so-called Wisconsin breast cancer data\n",
|
|
"using Logistic regression as our algorithm for classification. This is\n",
|
|
"a widely studied data set and can easily be included in demonstrations\n",
|
|
"of classification problems."
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "code",
|
|
"execution_count": 4,
|
|
"id": "9153234a",
|
|
"metadata": {
|
|
"collapsed": false,
|
|
"editable": true
|
|
},
|
|
"outputs": [],
|
|
"source": [
|
|
"%matplotlib inline\n",
|
|
"\n",
|
|
"import matplotlib.pyplot as plt\n",
|
|
"import numpy as np\n",
|
|
"from sklearn.model_selection import train_test_split \n",
|
|
"from sklearn.datasets import load_breast_cancer\n",
|
|
"from sklearn.linear_model import LogisticRegression\n",
|
|
"\n",
|
|
"# Load the data\n",
|
|
"cancer = load_breast_cancer()\n",
|
|
"\n",
|
|
"X_train, X_test, y_train, y_test = train_test_split(cancer.data,cancer.target,random_state=0)\n",
|
|
"print(X_train.shape)\n",
|
|
"print(X_test.shape)\n",
|
|
"# Logistic Regression\n",
|
|
"logreg = LogisticRegression(solver='lbfgs')\n",
|
|
"logreg.fit(X_train, y_train)\n",
|
|
"print(\"Test set accuracy with Logistic Regression: {:.2f}\".format(logreg.score(X_test,y_test)))"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "908d547b",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## Using the correlation matrix\n",
|
|
"\n",
|
|
"In addition to the above scores, we could also study the covariance (and the correlation matrix).\n",
|
|
"We use **Pandas** to compute the correlation matrix."
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "code",
|
|
"execution_count": 5,
|
|
"id": "8a46f4f3",
|
|
"metadata": {
|
|
"collapsed": false,
|
|
"editable": true
|
|
},
|
|
"outputs": [],
|
|
"source": [
|
|
"import matplotlib.pyplot as plt\n",
|
|
"import numpy as np\n",
|
|
"from sklearn.model_selection import train_test_split \n",
|
|
"from sklearn.datasets import load_breast_cancer\n",
|
|
"from sklearn.linear_model import LogisticRegression\n",
|
|
"cancer = load_breast_cancer()\n",
|
|
"import pandas as pd\n",
|
|
"# Making a data frame\n",
|
|
"cancerpd = pd.DataFrame(cancer.data, columns=cancer.feature_names)\n",
|
|
"\n",
|
|
"fig, axes = plt.subplots(15,2,figsize=(10,20))\n",
|
|
"malignant = cancer.data[cancer.target == 0]\n",
|
|
"benign = cancer.data[cancer.target == 1]\n",
|
|
"ax = axes.ravel()\n",
|
|
"\n",
|
|
"for i in range(30):\n",
|
|
" _, bins = np.histogram(cancer.data[:,i], bins =50)\n",
|
|
" ax[i].hist(malignant[:,i], bins = bins, alpha = 0.5)\n",
|
|
" ax[i].hist(benign[:,i], bins = bins, alpha = 0.5)\n",
|
|
" ax[i].set_title(cancer.feature_names[i])\n",
|
|
" ax[i].set_yticks(())\n",
|
|
"ax[0].set_xlabel(\"Feature magnitude\")\n",
|
|
"ax[0].set_ylabel(\"Frequency\")\n",
|
|
"ax[0].legend([\"Malignant\", \"Benign\"], loc =\"best\")\n",
|
|
"fig.tight_layout()\n",
|
|
"plt.show()\n",
|
|
"\n",
|
|
"import seaborn as sns\n",
|
|
"correlation_matrix = cancerpd.corr().round(1)\n",
|
|
"# use the heatmap function from seaborn to plot the correlation matrix\n",
|
|
"# annot = True to print the values inside the square\n",
|
|
"plt.figure(figsize=(15,8))\n",
|
|
"sns.heatmap(data=correlation_matrix, annot=True)\n",
|
|
"plt.show()"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "ba0275a7",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## Discussing the correlation data\n",
|
|
"\n",
|
|
"In the above example we note two things. In the first plot we display\n",
|
|
"the overlap of benign and malignant tumors as functions of the various\n",
|
|
"features in the Wisconsin data set. We see that for\n",
|
|
"some of the features we can distinguish clearly the benign and\n",
|
|
"malignant cases while for other features we cannot. This can point to\n",
|
|
"us which features may be of greater interest when we wish to classify\n",
|
|
"a benign or not benign tumour.\n",
|
|
"\n",
|
|
"In the second figure we have computed the so-called correlation\n",
|
|
"matrix, which in our case with thirty features becomes a $30\\times 30$\n",
|
|
"matrix.\n",
|
|
"\n",
|
|
"We constructed this matrix using **pandas** via the statements"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "code",
|
|
"execution_count": 6,
|
|
"id": "1af34f8e",
|
|
"metadata": {
|
|
"collapsed": false,
|
|
"editable": true
|
|
},
|
|
"outputs": [],
|
|
"source": [
|
|
"cancerpd = pd.DataFrame(cancer.data, columns=cancer.feature_names)"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "1eac30d3",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"and then"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "code",
|
|
"execution_count": 7,
|
|
"id": "a0cdd9c9",
|
|
"metadata": {
|
|
"collapsed": false,
|
|
"editable": true
|
|
},
|
|
"outputs": [],
|
|
"source": [
|
|
"correlation_matrix = cancerpd.corr().round(1)"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "013777ad",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"Diagonalizing this matrix we can in turn say something about which\n",
|
|
"features are of relevance and which are not. This leads us to\n",
|
|
"the classical Principal Component Analysis (PCA) theorem with\n",
|
|
"applications. This will be discussed later this semester."
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "410f90ac",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## Other measures in classification studies"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "code",
|
|
"execution_count": 8,
|
|
"id": "fa16a459",
|
|
"metadata": {
|
|
"collapsed": false,
|
|
"editable": true
|
|
},
|
|
"outputs": [],
|
|
"source": [
|
|
"import matplotlib.pyplot as plt\n",
|
|
"import numpy as np\n",
|
|
"from sklearn.model_selection import train_test_split \n",
|
|
"from sklearn.datasets import load_breast_cancer\n",
|
|
"from sklearn.linear_model import LogisticRegression\n",
|
|
"\n",
|
|
"# Load the data\n",
|
|
"cancer = load_breast_cancer()\n",
|
|
"\n",
|
|
"X_train, X_test, y_train, y_test = train_test_split(cancer.data,cancer.target,random_state=0)\n",
|
|
"print(X_train.shape)\n",
|
|
"print(X_test.shape)\n",
|
|
"# Logistic Regression\n",
|
|
"logreg = LogisticRegression(solver='lbfgs')\n",
|
|
"logreg.fit(X_train, y_train)\n",
|
|
"\n",
|
|
"from sklearn.preprocessing import LabelEncoder\n",
|
|
"from sklearn.model_selection import cross_validate\n",
|
|
"#Cross validation\n",
|
|
"accuracy = cross_validate(logreg,X_test,y_test,cv=10)['test_score']\n",
|
|
"print(accuracy)\n",
|
|
"print(\"Test set accuracy with Logistic Regression: {:.2f}\".format(logreg.score(X_test,y_test)))\n",
|
|
"\n",
|
|
"import scikitplot as skplt\n",
|
|
"y_pred = logreg.predict(X_test)\n",
|
|
"skplt.metrics.plot_confusion_matrix(y_test, y_pred, normalize=True)\n",
|
|
"plt.show()\n",
|
|
"y_probas = logreg.predict_proba(X_test)\n",
|
|
"skplt.metrics.plot_roc(y_test, y_probas)\n",
|
|
"plt.show()\n",
|
|
"skplt.metrics.plot_cumulative_gain(y_test, y_probas)\n",
|
|
"plt.show()"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "a721de53",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## Introduction to Neural networks\n",
|
|
"\n",
|
|
"Artificial neural networks are computational systems that can learn to\n",
|
|
"perform tasks by considering examples, generally without being\n",
|
|
"programmed with any task-specific rules. It is supposed to mimic a\n",
|
|
"biological system, wherein neurons interact by sending signals in the\n",
|
|
"form of mathematical functions between layers. All layers can contain\n",
|
|
"an arbitrary number of neurons, and each connection is represented by\n",
|
|
"a weight variable."
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "68de5052",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## Artificial neurons\n",
|
|
"\n",
|
|
"The field of artificial neural networks has a long history of\n",
|
|
"development, and is closely connected with the advancement of computer\n",
|
|
"science and computers in general. A model of artificial neurons was\n",
|
|
"first developed by McCulloch and Pitts in 1943 to study signal\n",
|
|
"processing in the brain and has later been refined by others. The\n",
|
|
"general idea is to mimic neural networks in the human brain, which is\n",
|
|
"composed of billions of neurons that communicate with each other by\n",
|
|
"sending electrical signals. Each neuron accumulates its incoming\n",
|
|
"signals, which must exceed an activation threshold to yield an\n",
|
|
"output. If the threshold is not overcome, the neuron remains inactive,\n",
|
|
"i.e. has zero output.\n",
|
|
"\n",
|
|
"This behaviour has inspired a simple mathematical model for an artificial neuron."
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "7685af02",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"<!-- Equation labels as ordinary links -->\n",
|
|
"<div id=\"artificialNeuron\"></div>\n",
|
|
"\n",
|
|
"$$\n",
|
|
"\\begin{equation}\n",
|
|
" y = f\\left(\\sum_{i=1}^n w_ix_i\\right) = f(u)\n",
|
|
"\\label{artificialNeuron} \\tag{1}\n",
|
|
"\\end{equation}\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "3dfcfcb0",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"Here, the output $y$ of the neuron is the value of its activation function, which have as input\n",
|
|
"a weighted sum of signals $x_i, \\dots ,x_n$ received by $n$ other neurons.\n",
|
|
"\n",
|
|
"Conceptually, it is helpful to divide neural networks into four\n",
|
|
"categories:\n",
|
|
"1. general purpose neural networks for supervised learning,\n",
|
|
"\n",
|
|
"2. neural networks designed specifically for image processing, the most prominent example of this class being Convolutional Neural Networks (CNNs),\n",
|
|
"\n",
|
|
"3. neural networks for sequential data such as Recurrent Neural Networks (RNNs), and\n",
|
|
"\n",
|
|
"4. neural networks for unsupervised learning such as Deep Boltzmann Machines.\n",
|
|
"\n",
|
|
"In natural science, DNNs and CNNs have already found numerous\n",
|
|
"applications. In statistical physics, they have been applied to detect\n",
|
|
"phase transitions in 2D Ising and Potts models, lattice gauge\n",
|
|
"theories, and different phases of polymers, or solving the\n",
|
|
"Navier-Stokes equation in weather forecasting. Deep learning has also\n",
|
|
"found interesting applications in quantum physics. Various quantum\n",
|
|
"phase transitions can be detected and studied using DNNs and CNNs,\n",
|
|
"topological phases, and even non-equilibrium many-body\n",
|
|
"localization. Representing quantum states as DNNs quantum state\n",
|
|
"tomography are among some of the impressive achievements to reveal the\n",
|
|
"potential of DNNs to facilitate the study of quantum systems.\n",
|
|
"\n",
|
|
"In quantum information theory, it has been shown that one can perform\n",
|
|
"gate decompositions with the help of neural. \n",
|
|
"\n",
|
|
"The applications are not limited to the natural sciences. There is a\n",
|
|
"plethora of applications in essentially all disciplines, from the\n",
|
|
"humanities to life science and medicine."
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "0d037ca7",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## Neural network types\n",
|
|
"\n",
|
|
"An artificial neural network (ANN), is a computational model that\n",
|
|
"consists of layers of connected neurons, or nodes or units. We will\n",
|
|
"refer to these interchangeably as units or nodes, and sometimes as\n",
|
|
"neurons.\n",
|
|
"\n",
|
|
"It is supposed to mimic a biological nervous system by letting each\n",
|
|
"neuron interact with other neurons by sending signals in the form of\n",
|
|
"mathematical functions between layers. A wide variety of different\n",
|
|
"ANNs have been developed, but most of them consist of an input layer,\n",
|
|
"an output layer and eventual layers in-between, called *hidden\n",
|
|
"layers*. All layers can contain an arbitrary number of nodes, and each\n",
|
|
"connection between two nodes is associated with a weight variable.\n",
|
|
"\n",
|
|
"Neural networks (also called neural nets) are neural-inspired\n",
|
|
"nonlinear models for supervised learning. As we will see, neural nets\n",
|
|
"can be viewed as natural, more powerful extensions of supervised\n",
|
|
"learning methods such as linear and logistic regression and soft-max\n",
|
|
"methods we discussed earlier."
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "7bcf7188",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## Feed-forward neural networks\n",
|
|
"\n",
|
|
"The feed-forward neural network (FFNN) was the first and simplest type\n",
|
|
"of ANNs that were devised. In this network, the information moves in\n",
|
|
"only one direction: forward through the layers.\n",
|
|
"\n",
|
|
"Nodes are represented by circles, while the arrows display the\n",
|
|
"connections between the nodes, including the direction of information\n",
|
|
"flow. Additionally, each arrow corresponds to a weight variable\n",
|
|
"(figure to come). We observe that each node in a layer is connected\n",
|
|
"to *all* nodes in the subsequent layer, making this a so-called\n",
|
|
"*fully-connected* FFNN."
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "cd094e20",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## Convolutional Neural Network\n",
|
|
"\n",
|
|
"A different variant of FFNNs are *convolutional neural networks*\n",
|
|
"(CNNs), which have a connectivity pattern inspired by the animal\n",
|
|
"visual cortex. Individual neurons in the visual cortex only respond to\n",
|
|
"stimuli from small sub-regions of the visual field, called a receptive\n",
|
|
"field. This makes the neurons well-suited to exploit the strong\n",
|
|
"spatially local correlation present in natural images. The response of\n",
|
|
"each neuron can be approximated mathematically as a convolution\n",
|
|
"operation. (figure to come)\n",
|
|
"\n",
|
|
"Convolutional neural networks emulate the behaviour of neurons in the\n",
|
|
"visual cortex by enforcing a *local* connectivity pattern between\n",
|
|
"nodes of adjacent layers: Each node in a convolutional layer is\n",
|
|
"connected only to a subset of the nodes in the previous layer, in\n",
|
|
"contrast to the fully-connected FFNN. Often, CNNs consist of several\n",
|
|
"convolutional layers that learn local features of the input, with a\n",
|
|
"fully-connected layer at the end, which gathers all the local data and\n",
|
|
"produces the outputs. They have wide applications in image and video\n",
|
|
"recognition."
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "ea99157e",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## Recurrent neural networks\n",
|
|
"\n",
|
|
"So far we have only mentioned ANNs where information flows in one\n",
|
|
"direction: forward. *Recurrent neural networks* on the other hand,\n",
|
|
"have connections between nodes that form directed *cycles*. This\n",
|
|
"creates a form of internal memory which are able to capture\n",
|
|
"information on what has been calculated before; the output is\n",
|
|
"dependent on the previous computations. Recurrent NNs make use of\n",
|
|
"sequential information by performing the same task for every element\n",
|
|
"in a sequence, where each element depends on previous elements. An\n",
|
|
"example of such information is sentences, making recurrent NNs\n",
|
|
"especially well-suited for handwriting and speech recognition."
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "b73754c2",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## Other types of networks\n",
|
|
"\n",
|
|
"There are many other kinds of ANNs that have been developed. One type\n",
|
|
"that is specifically designed for interpolation in multidimensional\n",
|
|
"space is the radial basis function (RBF) network. RBFs are typically\n",
|
|
"made up of three layers: an input layer, a hidden layer with\n",
|
|
"non-linear radial symmetric activation functions and a linear output\n",
|
|
"layer (''linear'' here means that each node in the output layer has a\n",
|
|
"linear activation function). The layers are normally fully-connected\n",
|
|
"and there are no cycles, thus RBFs can be viewed as a type of\n",
|
|
"fully-connected FFNN. They are however usually treated as a separate\n",
|
|
"type of NN due the unusual activation functions."
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "aa97c83d",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## Multilayer perceptrons\n",
|
|
"\n",
|
|
"One uses often so-called fully-connected feed-forward neural networks\n",
|
|
"with three or more layers (an input layer, one or more hidden layers\n",
|
|
"and an output layer) consisting of neurons that have non-linear\n",
|
|
"activation functions.\n",
|
|
"\n",
|
|
"Such networks are often called *multilayer perceptrons* (MLPs)."
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "abe84919",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## Why multilayer perceptrons?\n",
|
|
"\n",
|
|
"According to the *Universal approximation theorem*, a feed-forward\n",
|
|
"neural network with just a single hidden layer containing a finite\n",
|
|
"number of neurons can approximate a continuous multidimensional\n",
|
|
"function to arbitrary accuracy, assuming the activation function for\n",
|
|
"the hidden layer is a **non-constant, bounded and\n",
|
|
"monotonically-increasing continuous function**.\n",
|
|
"\n",
|
|
"Note that the requirements on the activation function only applies to\n",
|
|
"the hidden layer, the output nodes are always assumed to be linear, so\n",
|
|
"as to not restrict the range of output values."
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "d3ff207b",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## Illustration of a single perceptron model and a multi-perceptron model\n",
|
|
"\n",
|
|
"<!-- dom:FIGURE: [figures/nns.png, width=600 frac=0.8] In a) we show a single perceptron model while in b) we dispay a network with two hidden layers, an input layer and an output layer. -->\n",
|
|
"<!-- begin figure -->\n",
|
|
"\n",
|
|
"<img src=\"figures/nns.png\" width=\"600\"><p style=\"font-size: 0.9em\"><i>Figure 1: In a) we show a single perceptron model while in b) we dispay a network with two hidden layers, an input layer and an output layer.</i></p>\n",
|
|
"<!-- end figure -->"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "f982c11f",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## Examples of XOR, OR and AND gates\n",
|
|
"\n",
|
|
"Let us first try to fit various gates using standard linear\n",
|
|
"regression. The gates we are thinking of are the classical XOR, OR and\n",
|
|
"AND gates, well-known elements in computer science. The tables here\n",
|
|
"show how we can set up the inputs $x_1$ and $x_2$ in order to yield a\n",
|
|
"specific target $y_i$."
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "code",
|
|
"execution_count": 9,
|
|
"id": "04a3e090",
|
|
"metadata": {
|
|
"collapsed": false,
|
|
"editable": true
|
|
},
|
|
"outputs": [],
|
|
"source": [
|
|
"\"\"\"\n",
|
|
"Simple code that tests XOR, OR and AND gates with linear regression\n",
|
|
"\"\"\"\n",
|
|
"\n",
|
|
"import numpy as np\n",
|
|
"# Design matrix\n",
|
|
"X = np.array([ [1, 0, 0], [1, 0, 1], [1, 1, 0],[1, 1, 1]],dtype=np.float64)\n",
|
|
"print(f\"The X.TX matrix:{X.T @ X}\")\n",
|
|
"Xinv = np.linalg.pinv(X.T @ X)\n",
|
|
"print(f\"The invers of X.TX matrix:{Xinv}\")\n",
|
|
"\n",
|
|
"# The XOR gate \n",
|
|
"yXOR = np.array( [ 0, 1 ,1, 0])\n",
|
|
"ThetaXOR = Xinv @ X.T @ yXOR\n",
|
|
"print(f\"The values of theta for the XOR gate:{ThetaXOR}\")\n",
|
|
"print(f\"The linear regression prediction for the XOR gate:{X @ ThetaXOR}\")\n",
|
|
"\n",
|
|
"\n",
|
|
"# The OR gate \n",
|
|
"yOR = np.array( [ 0, 1 ,1, 1])\n",
|
|
"ThetaOR = Xinv @ X.T @ yOR\n",
|
|
"print(f\"The values of theta for the OR gate:{ThetaOR}\")\n",
|
|
"print(f\"The linear regression prediction for the OR gate:{X @ ThetaOR}\")\n",
|
|
"\n",
|
|
"\n",
|
|
"# The OR gate \n",
|
|
"yAND = np.array( [ 0, 0 ,0, 1])\n",
|
|
"ThetaAND = Xinv @ X.T @ yAND\n",
|
|
"print(f\"The values of theta for the AND gate:{ThetaAND}\")\n",
|
|
"print(f\"The linear regression prediction for the AND gate:{X @ ThetaAND}\")"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "95b1f5a5",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"What is happening here?"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "0d200eff",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## Does Logistic Regression do a better Job?"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "code",
|
|
"execution_count": 10,
|
|
"id": "040a69d0",
|
|
"metadata": {
|
|
"collapsed": false,
|
|
"editable": true
|
|
},
|
|
"outputs": [],
|
|
"source": [
|
|
"\"\"\"\n",
|
|
"Simple code that tests XOR and OR gates with linear regression\n",
|
|
"and logistic regression\n",
|
|
"\"\"\"\n",
|
|
"\n",
|
|
"import matplotlib.pyplot as plt\n",
|
|
"from sklearn.linear_model import LogisticRegression\n",
|
|
"import numpy as np\n",
|
|
"\n",
|
|
"# Design matrix\n",
|
|
"X = np.array([ [1, 0, 0], [1, 0, 1], [1, 1, 0],[1, 1, 1]],dtype=np.float64)\n",
|
|
"print(f\"The X.TX matrix:{X.T @ X}\")\n",
|
|
"Xinv = np.linalg.pinv(X.T @ X)\n",
|
|
"print(f\"The invers of X.TX matrix:{Xinv}\")\n",
|
|
"\n",
|
|
"# The XOR gate \n",
|
|
"yXOR = np.array( [ 0, 1 ,1, 0])\n",
|
|
"ThetaXOR = Xinv @ X.T @ yXOR\n",
|
|
"print(f\"The values of theta for the XOR gate:{ThetaXOR}\")\n",
|
|
"print(f\"The linear regression prediction for the XOR gate:{X @ ThetaXOR}\")\n",
|
|
"\n",
|
|
"\n",
|
|
"# The OR gate \n",
|
|
"yOR = np.array( [ 0, 1 ,1, 1])\n",
|
|
"ThetaOR = Xinv @ X.T @ yOR\n",
|
|
"print(f\"The values of theta for the OR gate:{ThetaOR}\")\n",
|
|
"print(f\"The linear regression prediction for the OR gate:{X @ ThetaOR}\")\n",
|
|
"\n",
|
|
"\n",
|
|
"# The OR gate \n",
|
|
"yAND = np.array( [ 0, 0 ,0, 1])\n",
|
|
"ThetaAND = Xinv @ X.T @ yAND\n",
|
|
"print(f\"The values of theta for the AND gate:{ThetaAND}\")\n",
|
|
"print(f\"The linear regression prediction for the AND gate:{X @ ThetaAND}\")\n",
|
|
"\n",
|
|
"# Now we change to logistic regression\n",
|
|
"\n",
|
|
"\n",
|
|
"# Logistic Regression\n",
|
|
"logreg = LogisticRegression()\n",
|
|
"logreg.fit(X, yOR)\n",
|
|
"print(\"Test set accuracy with Logistic Regression for OR gate: {:.2f}\".format(logreg.score(X,yOR)))\n",
|
|
"\n",
|
|
"logreg.fit(X, yXOR)\n",
|
|
"print(\"Test set accuracy with Logistic Regression for XOR gate: {:.2f}\".format(logreg.score(X,yXOR)))\n",
|
|
"\n",
|
|
"\n",
|
|
"logreg.fit(X, yAND)\n",
|
|
"print(\"Test set accuracy with Logistic Regression for AND gate: {:.2f}\".format(logreg.score(X,yAND)))"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "49f17f65",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"Not exactly impressive, but somewhat better."
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "714e0891",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## Adding Neural Networks"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "code",
|
|
"execution_count": 11,
|
|
"id": "28bde670",
|
|
"metadata": {
|
|
"collapsed": false,
|
|
"editable": true
|
|
},
|
|
"outputs": [],
|
|
"source": [
|
|
"\n",
|
|
"# and now neural networks with Scikit-Learn and the XOR\n",
|
|
"\n",
|
|
"from sklearn.neural_network import MLPClassifier\n",
|
|
"from sklearn.datasets import make_classification\n",
|
|
"X, yXOR = make_classification(n_samples=100, random_state=1)\n",
|
|
"FFNN = MLPClassifier(random_state=1, max_iter=300).fit(X, yXOR)\n",
|
|
"FFNN.predict_proba(X)\n",
|
|
"print(f\"Test set accuracy with Feed Forward Neural Network for XOR gate:{FFNN.score(X, yXOR)}\")"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "4440856f",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## Mathematical model\n",
|
|
"\n",
|
|
"The output $y$ is produced via the activation function $f$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "6199da92",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"$$\n",
|
|
"y = f\\left(\\sum_{i=1}^n w_ix_i + b_i\\right) = f(z),\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "62c964e3",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"This function receives $x_i$ as inputs.\n",
|
|
"Here the activation $z=(\\sum_{i=1}^n w_ix_i+b_i)$. \n",
|
|
"In an FFNN of such neurons, the *inputs* $x_i$ are the *outputs* of\n",
|
|
"the neurons in the preceding layer. Furthermore, an MLP is\n",
|
|
"fully-connected, which means that each neuron receives a weighted sum\n",
|
|
"of the outputs of *all* neurons in the previous layer."
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "64ba4c70",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## Mathematical model\n",
|
|
"\n",
|
|
"First, for each node $i$ in the first hidden layer, we calculate a weighted sum $z_i^1$ of the input coordinates $x_j$,"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "66c11135",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"<!-- Equation labels as ordinary links -->\n",
|
|
"<div id=\"_auto1\"></div>\n",
|
|
"\n",
|
|
"$$\n",
|
|
"\\begin{equation} z_i^1 = \\sum_{j=1}^{M} w_{ij}^1 x_j + b_i^1\n",
|
|
"\\label{_auto1} \\tag{2}\n",
|
|
"\\end{equation}\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "0f47b20a",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"Here $b_i$ is the so-called bias which is normally needed in\n",
|
|
"case of zero activation weights or inputs. How to fix the biases and\n",
|
|
"the weights will be discussed below. The value of $z_i^1$ is the\n",
|
|
"argument to the activation function $f_i$ of each node $i$, The\n",
|
|
"variable $M$ stands for all possible inputs to a given node $i$ in the\n",
|
|
"first layer. We define the output $y_i^1$ of all neurons in layer 1 as"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "bda56156",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"<!-- Equation labels as ordinary links -->\n",
|
|
"<div id=\"outputLayer1\"></div>\n",
|
|
"\n",
|
|
"$$\n",
|
|
"\\begin{equation}\n",
|
|
" y_i^1 = f(z_i^1) = f\\left(\\sum_{j=1}^M w_{ij}^1 x_j + b_i^1\\right)\n",
|
|
"\\label{outputLayer1} \\tag{3}\n",
|
|
"\\end{equation}\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "1330fab9",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"where we assume that all nodes in the same layer have identical\n",
|
|
"activation functions, hence the notation $f$. In general, we could assume in the more general case that different layers have different activation functions.\n",
|
|
"In this case we would identify these functions with a superscript $l$ for the $l$-th layer,"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "ae474dfb",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"<!-- Equation labels as ordinary links -->\n",
|
|
"<div id=\"generalLayer\"></div>\n",
|
|
"\n",
|
|
"$$\n",
|
|
"\\begin{equation}\n",
|
|
" y_i^l = f^l(u_i^l) = f^l\\left(\\sum_{j=1}^{N_{l-1}} w_{ij}^l y_j^{l-1} + b_i^l\\right)\n",
|
|
"\\label{generalLayer} \\tag{4}\n",
|
|
"\\end{equation}\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "b6cb6fed",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"where $N_l$ is the number of nodes in layer $l$. When the output of\n",
|
|
"all the nodes in the first hidden layer are computed, the values of\n",
|
|
"the subsequent layer can be calculated and so forth until the output\n",
|
|
"is obtained."
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "2f8f9b4e",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## Mathematical model\n",
|
|
"\n",
|
|
"The output of neuron $i$ in layer 2 is thus,"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "18e74238",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"<!-- Equation labels as ordinary links -->\n",
|
|
"<div id=\"_auto2\"></div>\n",
|
|
"\n",
|
|
"$$\n",
|
|
"\\begin{equation}\n",
|
|
" y_i^2 = f^2\\left(\\sum_{j=1}^N w_{ij}^2 y_j^1 + b_i^2\\right) \n",
|
|
"\\label{_auto2} \\tag{5}\n",
|
|
"\\end{equation}\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "d10df3e7",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"<!-- Equation labels as ordinary links -->\n",
|
|
"<div id=\"outputLayer2\"></div>\n",
|
|
"\n",
|
|
"$$\n",
|
|
"\\begin{equation} \n",
|
|
" = f^2\\left[\\sum_{j=1}^N w_{ij}^2f^1\\left(\\sum_{k=1}^M w_{jk}^1 x_k + b_j^1\\right) + b_i^2\\right]\n",
|
|
"\\label{outputLayer2} \\tag{6}\n",
|
|
"\\end{equation}\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "da21a316",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"where we have substituted $y_k^1$ with the inputs $x_k$. Finally, the ANN output reads"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "76938a28",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"<!-- Equation labels as ordinary links -->\n",
|
|
"<div id=\"_auto3\"></div>\n",
|
|
"\n",
|
|
"$$\n",
|
|
"\\begin{equation}\n",
|
|
" y_i^3 = f^3\\left(\\sum_{j=1}^N w_{ij}^3 y_j^2 + b_i^3\\right) \n",
|
|
"\\label{_auto3} \\tag{7}\n",
|
|
"\\end{equation}\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "65434967",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"<!-- Equation labels as ordinary links -->\n",
|
|
"<div id=\"_auto4\"></div>\n",
|
|
"\n",
|
|
"$$\n",
|
|
"\\begin{equation} \n",
|
|
" = f_3\\left[\\sum_{j} w_{ij}^3 f^2\\left(\\sum_{k} w_{jk}^2 f^1\\left(\\sum_{m} w_{km}^1 x_m + b_k^1\\right) + b_j^2\\right)\n",
|
|
" + b_1^3\\right]\n",
|
|
"\\label{_auto4} \\tag{8}\n",
|
|
"\\end{equation}\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "31d4f5aa",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## Mathematical model\n",
|
|
"\n",
|
|
"We can generalize this expression to an MLP with $l$ hidden\n",
|
|
"layers. The complete functional form is,"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "114030e5",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"<!-- Equation labels as ordinary links -->\n",
|
|
"<div id=\"completeNN\"></div>\n",
|
|
"\n",
|
|
"$$\n",
|
|
"\\begin{equation}\n",
|
|
"y^{l+1}_i = f^{l+1}\\left[\\!\\sum_{j=1}^{N_l} w_{ij}^3 f^l\\left(\\sum_{k=1}^{N_{l-1}}w_{jk}^{l-1}\\left(\\dots f^1\\left(\\sum_{n=1}^{N_0} w_{mn}^1 x_n+ b_m^1\\right)\\dots\\right)+b_k^2\\right)+b_1^3\\right] \n",
|
|
"\\label{completeNN} \\tag{9}\n",
|
|
"\\end{equation}\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "a93aec4e",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"which illustrates a basic property of MLPs: The only independent\n",
|
|
"variables are the input values $x_n$."
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "7c85562d",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"## Mathematical model\n",
|
|
"\n",
|
|
"This confirms that an MLP, despite its quite convoluted mathematical\n",
|
|
"form, is nothing more than an analytic function, specifically a\n",
|
|
"mapping of real-valued vectors $\\hat{x} \\in \\mathbb{R}^n \\rightarrow\n",
|
|
"\\hat{y} \\in \\mathbb{R}^m$.\n",
|
|
"\n",
|
|
"Furthermore, the flexibility and universality of an MLP can be\n",
|
|
"illustrated by realizing that the expression is essentially a nested\n",
|
|
"sum of scaled activation functions of the form"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "1152ea5e",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"<!-- Equation labels as ordinary links -->\n",
|
|
"<div id=\"_auto5\"></div>\n",
|
|
"\n",
|
|
"$$\n",
|
|
"\\begin{equation}\n",
|
|
" f(x) = c_1 f(c_2 x + c_3) + c_4\n",
|
|
"\\label{_auto5} \\tag{10}\n",
|
|
"\\end{equation}\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "4f3d4b33",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"where the parameters $c_i$ are weights and biases. By adjusting these\n",
|
|
"parameters, the activation functions can be shifted up and down or\n",
|
|
"left and right, change slope or be rescaled which is the key to the\n",
|
|
"flexibility of a neural network."
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "4c1ac54e",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"### Matrix-vector notation\n",
|
|
"\n",
|
|
"We can introduce a more convenient notation for the activations in an A NN. \n",
|
|
"\n",
|
|
"Additionally, we can represent the biases and activations\n",
|
|
"as layer-wise column vectors $\\hat{b}_l$ and $\\hat{y}_l$, so that the $i$-th element of each vector \n",
|
|
"is the bias $b_i^l$ and activation $y_i^l$ of node $i$ in layer $l$ respectively. \n",
|
|
"\n",
|
|
"We have that $\\mathrm{W}_l$ is an $N_{l-1} \\times N_l$ matrix, while $\\hat{b}_l$ and $\\hat{y}_l$ are $N_l \\times 1$ column vectors. \n",
|
|
"With this notation, the sum becomes a matrix-vector multiplication, and we can write\n",
|
|
"the equation for the activations of hidden layer 2 (assuming three nodes for simplicity) as"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "5c4a861f",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"<!-- Equation labels as ordinary links -->\n",
|
|
"<div id=\"_auto6\"></div>\n",
|
|
"\n",
|
|
"$$\n",
|
|
"\\begin{equation}\n",
|
|
" \\hat{y}_2 = f_2(\\mathrm{W}_2 \\hat{y}_{1} + \\hat{b}_{2}) = \n",
|
|
" f_2\\left(\\left[\\begin{array}{ccc}\n",
|
|
" w^2_{11} &w^2_{12} &w^2_{13} \\\\\n",
|
|
" w^2_{21} &w^2_{22} &w^2_{23} \\\\\n",
|
|
" w^2_{31} &w^2_{32} &w^2_{33} \\\\\n",
|
|
" \\end{array} \\right] \\cdot\n",
|
|
" \\left[\\begin{array}{c}\n",
|
|
" y^1_1 \\\\\n",
|
|
" y^1_2 \\\\\n",
|
|
" y^1_3 \\\\\n",
|
|
" \\end{array}\\right] + \n",
|
|
" \\left[\\begin{array}{c}\n",
|
|
" b^2_1 \\\\\n",
|
|
" b^2_2 \\\\\n",
|
|
" b^2_3 \\\\\n",
|
|
" \\end{array}\\right]\\right).\n",
|
|
"\\label{_auto6} \\tag{11}\n",
|
|
"\\end{equation}\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "276b271b",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"### Matrix-vector notation and activation\n",
|
|
"\n",
|
|
"The activation of node $i$ in layer 2 is"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "63a5b8f1",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"<!-- Equation labels as ordinary links -->\n",
|
|
"<div id=\"_auto7\"></div>\n",
|
|
"\n",
|
|
"$$\n",
|
|
"\\begin{equation}\n",
|
|
" y^2_i = f_2\\Bigr(w^2_{i1}y^1_1 + w^2_{i2}y^1_2 + w^2_{i3}y^1_3 + b^2_i\\Bigr) = \n",
|
|
" f_2\\left(\\sum_{j=1}^3 w^2_{ij} y_j^1 + b^2_i\\right).\n",
|
|
"\\label{_auto7} \\tag{12}\n",
|
|
"\\end{equation}\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "316b8c32",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"This is not just a convenient and compact notation, but also a useful\n",
|
|
"and intuitive way to think about MLPs: The output is calculated by a\n",
|
|
"series of matrix-vector multiplications and vector additions that are\n",
|
|
"used as input to the activation functions. For each operation\n",
|
|
"$\\mathrm{W}_l \\hat{y}_{l-1}$ we move forward one layer."
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "34ba90c8",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"### Activation functions\n",
|
|
"\n",
|
|
"A property that characterizes a neural network, other than its\n",
|
|
"connectivity, is the choice of activation function(s). As described\n",
|
|
"in, the following restrictions are imposed on an activation function\n",
|
|
"for a FFNN to fulfill the universal approximation theorem\n",
|
|
"\n",
|
|
" * Non-constant\n",
|
|
"\n",
|
|
" * Bounded\n",
|
|
"\n",
|
|
" * Monotonically-increasing\n",
|
|
"\n",
|
|
" * Continuous"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "3019fcaf",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"### Activation functions, Logistic and Hyperbolic ones\n",
|
|
"\n",
|
|
"The second requirement excludes all linear functions. Furthermore, in\n",
|
|
"a MLP with only linear activation functions, each layer simply\n",
|
|
"performs a linear transformation of its inputs.\n",
|
|
"\n",
|
|
"Regardless of the number of layers, the output of the NN will be\n",
|
|
"nothing but a linear function of the inputs. Thus we need to introduce\n",
|
|
"some kind of non-linearity to the NN to be able to fit non-linear\n",
|
|
"functions Typical examples are the logistic *Sigmoid*"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "389ff36b",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"$$\n",
|
|
"f(x) = \\frac{1}{1 + e^{-x}},\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "ee9b399a",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"and the *hyperbolic tangent* function"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "36f98b26",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"$$\n",
|
|
"f(x) = \\tanh(x)\n",
|
|
"$$"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"id": "cb7b8839",
|
|
"metadata": {
|
|
"editable": true
|
|
},
|
|
"source": [
|
|
"### Relevance\n",
|
|
"\n",
|
|
"The *sigmoid* function are more biologically plausible because the\n",
|
|
"output of inactive neurons are zero. Such activation function are\n",
|
|
"called *one-sided*. However, it has been shown that the hyperbolic\n",
|
|
"tangent performs better than the sigmoid for training MLPs. has\n",
|
|
"become the most popular for *deep neural networks*"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "code",
|
|
"execution_count": 12,
|
|
"id": "db8d28b5",
|
|
"metadata": {
|
|
"collapsed": false,
|
|
"editable": true
|
|
},
|
|
"outputs": [],
|
|
"source": [
|
|
"\"\"\"The sigmoid function (or the logistic curve) is a \n",
|
|
"function that takes any real number, z, and outputs a number (0,1).\n",
|
|
"It is useful in neural networks for assigning weights on a relative scale.\n",
|
|
"The value z is the weighted sum of parameters involved in the learning algorithm.\"\"\"\n",
|
|
"\n",
|
|
"import numpy\n",
|
|
"import matplotlib.pyplot as plt\n",
|
|
"import math as mt\n",
|
|
"\n",
|
|
"z = numpy.arange(-5, 5, .1)\n",
|
|
"sigma_fn = numpy.vectorize(lambda z: 1/(1+numpy.exp(-z)))\n",
|
|
"sigma = sigma_fn(z)\n",
|
|
"\n",
|
|
"fig = plt.figure()\n",
|
|
"ax = fig.add_subplot(111)\n",
|
|
"ax.plot(z, sigma)\n",
|
|
"ax.set_ylim([-0.1, 1.1])\n",
|
|
"ax.set_xlim([-5,5])\n",
|
|
"ax.grid(True)\n",
|
|
"ax.set_xlabel('z')\n",
|
|
"ax.set_title('sigmoid function')\n",
|
|
"\n",
|
|
"plt.show()\n",
|
|
"\n",
|
|
"\"\"\"Step Function\"\"\"\n",
|
|
"z = numpy.arange(-5, 5, .02)\n",
|
|
"step_fn = numpy.vectorize(lambda z: 1.0 if z >= 0.0 else 0.0)\n",
|
|
"step = step_fn(z)\n",
|
|
"\n",
|
|
"fig = plt.figure()\n",
|
|
"ax = fig.add_subplot(111)\n",
|
|
"ax.plot(z, step)\n",
|
|
"ax.set_ylim([-0.5, 1.5])\n",
|
|
"ax.set_xlim([-5,5])\n",
|
|
"ax.grid(True)\n",
|
|
"ax.set_xlabel('z')\n",
|
|
"ax.set_title('step function')\n",
|
|
"\n",
|
|
"plt.show()\n",
|
|
"\n",
|
|
"\"\"\"Sine Function\"\"\"\n",
|
|
"z = numpy.arange(-2*mt.pi, 2*mt.pi, 0.1)\n",
|
|
"t = numpy.sin(z)\n",
|
|
"\n",
|
|
"fig = plt.figure()\n",
|
|
"ax = fig.add_subplot(111)\n",
|
|
"ax.plot(z, t)\n",
|
|
"ax.set_ylim([-1.0, 1.0])\n",
|
|
"ax.set_xlim([-2*mt.pi,2*mt.pi])\n",
|
|
"ax.grid(True)\n",
|
|
"ax.set_xlabel('z')\n",
|
|
"ax.set_title('sine function')\n",
|
|
"\n",
|
|
"plt.show()\n",
|
|
"\n",
|
|
"\"\"\"Plots a graph of the squashing function used by a rectified linear\n",
|
|
"unit\"\"\"\n",
|
|
"z = numpy.arange(-2, 2, .1)\n",
|
|
"zero = numpy.zeros(len(z))\n",
|
|
"y = numpy.max([zero, z], axis=0)\n",
|
|
"\n",
|
|
"fig = plt.figure()\n",
|
|
"ax = fig.add_subplot(111)\n",
|
|
"ax.plot(z, y)\n",
|
|
"ax.set_ylim([-2.0, 2.0])\n",
|
|
"ax.set_xlim([-2.0, 2.0])\n",
|
|
"ax.grid(True)\n",
|
|
"ax.set_xlabel('z')\n",
|
|
"ax.set_title('Rectified linear unit')\n",
|
|
"\n",
|
|
"plt.show()"
|
|
]
|
|
}
|
|
],
|
|
"metadata": {},
|
|
"nbformat": 4,
|
|
"nbformat_minor": 5
|
|
}
|