1476 lines
47 KiB
Plaintext
1476 lines
47 KiB
Plaintext
{
|
||
"cells": [
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"<!-- dom:TITLE: Data Analysis and Machine Learning Lectures: Optimization and Gradient Methods -->\n",
|
||
"# Data Analysis and Machine Learning Lectures: Optimization and Gradient Methods\n",
|
||
"<!-- dom:AUTHOR: Morten Hjorth-Jensen at Department of Physics, University of Oslo & Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University -->\n",
|
||
"<!-- Author: --> \n",
|
||
"**Morten Hjorth-Jensen**, Department of Physics, University of Oslo and Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University\n",
|
||
"\n",
|
||
"Date: **Sep 27, 2018**\n",
|
||
"\n",
|
||
"Copyright 1999-2018, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license\n",
|
||
"\n",
|
||
"\n",
|
||
"\n",
|
||
"\n",
|
||
"## Optimization, the central part of any Machine Learning algortithm\n",
|
||
"\n",
|
||
"Almost every problem in machine learning and data science starts with\n",
|
||
"a dataset $X$, a model $g(\\beta)$, which is a function of the\n",
|
||
"parameters $\\beta$ and a cost function $C(X, g(\\beta))$ that allows\n",
|
||
"us to judge how well the model $g(\\beta)$ explains the observations\n",
|
||
"$X$. The model is fit by finding the values of $\\beta$ that minimize\n",
|
||
"the cost function. Ideally we would be able to solve for $\\beta$\n",
|
||
"analytically, however this is not possible in general and we must use\n",
|
||
"some approximative/numerical method to compute the minimum.\n",
|
||
"\n",
|
||
"\n",
|
||
"## Revisiting our Logistic Regression case\n",
|
||
"\n",
|
||
"In our discussion on Logistic Regression we defined we studied first the \n",
|
||
"case of\n",
|
||
"two classes, with $y_i$ either\n",
|
||
"$0$ or $1$. Furthermore we assumed also that we have only two\n",
|
||
"parameters $\\beta$ in our fitting of the Sigmoid function, that is we\n",
|
||
"defined probabilities"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\begin{align*}\n",
|
||
"p(y_i=1|x_i,\\hat{\\beta}) &= \\frac{\\exp{(\\beta_0+\\beta_1x_i)}}{1+\\exp{(\\beta_0+\\beta_1x_i)}},\\nonumber\\\\\n",
|
||
"p(y_i=0|x_i,\\hat{\\beta}) &= 1 - p(y_i=1|x_i,\\hat{\\beta}),\n",
|
||
"\\end{align*}\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"where $\\hat{\\beta}$ are the weights we wish to extract from data, in our case $\\beta_0$ and $\\beta_1$. \n",
|
||
"\n",
|
||
"## The equations to solve\n",
|
||
"\n",
|
||
"Our compact equations used a definition of a vector $\\hat{y}$ with $n$\n",
|
||
"elements $y_i$, an $n\\times p$ matrix $\\hat{X}$ which contains the\n",
|
||
"$x_i$ values and a vector $\\hat{p}$ of fitted probabilities\n",
|
||
"$p(y_i\\vert x_i,\\hat{\\beta})$. We rewrote in a more compact form\n",
|
||
"the first derivative of the cost function as"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\frac{\\partial \\mathcal{C}(\\hat{\\beta})}{\\partial \\hat{\\beta}} = -\\hat{X}^T\\left(\\hat{y}-\\hat{p}\\right).\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"If we in addition define a diagonal matrix $\\hat{W}$ with elements \n",
|
||
"$p(y_i\\vert x_i,\\hat{\\beta})(1-p(y_i\\vert x_i,\\hat{\\beta})$, we can obtain a compact expression of the second derivative as"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\frac{\\partial^2 \\mathcal{C}(\\hat{\\beta})}{\\partial \\hat{\\beta}\\partial \\hat{\\beta}^T} = \\hat{X}^T\\hat{W}\\hat{X}.\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"This defines what we call the Hessian.\n",
|
||
"\n",
|
||
"## Solving using Newton-Raphson's method\n",
|
||
"\n",
|
||
"If we can set up these equations, Newton-Raphson's iterative method is the nomrally the method of choice. It requires however that we setting the matrices that define the first and second derivatives. \n",
|
||
"\n",
|
||
"Our iterative scheme is then given by"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{\\beta}^{\\mathrm{new}} = \\hat{\\beta}^{\\mathrm{old}}-\\left(\\frac{\\partial^2 \\mathcal{C}(\\hat{\\beta})}{\\partial \\hat{\\beta}\\partial \\hat{\\beta}^T}\\right)^{-1}\\times \\left(\\frac{\\partial \\mathcal{C}(\\hat{\\beta})}{\\partial \\hat{\\beta}}\\right)_{\\hat{\\beta}^{\\mathrm{old}}},\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"or in matrix form as"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{\\beta}^{\\mathrm{new}} = \\hat{\\beta}^{\\mathrm{old}}-\\left(\\hat{X}^T\\hat{W}\\hat{X} \\right)^{-1}\\times \\left(-\\hat{X}^T(\\hat{y}-\\hat{p}) \\right)_{\\hat{\\beta}^{\\mathrm{old}}}.\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"The right-hand side is computed with the old values of $\\beta$. \n",
|
||
"\n",
|
||
"If we can compute these matrices, in particular the Hessian, the above is often the easiest method to implement. \n",
|
||
"\n",
|
||
"\n",
|
||
"## Brief reminder on Newton-Raphson's method\n",
|
||
"\n",
|
||
"Let us quicly remind ourselves how we derive the above method.\n",
|
||
"\n",
|
||
"Perhaps the most celebrated of all one-dimensional root-finding\n",
|
||
"routines is Newton's method, also called the Newton-Raphson\n",
|
||
"method. This method is distinguished from the previously discussed\n",
|
||
"methods by the fact that it requires the evaluation of both the\n",
|
||
"function $f$ and its derivative $f'$ at arbitrary points. In this\n",
|
||
"sense, it is taylored to cases with e.g., transcendental equations.\n",
|
||
"If you can only calculate the derivative\n",
|
||
"numerically and/or your function is not of the smooth type, we\n",
|
||
"discourage the use of this method.\n",
|
||
"\n",
|
||
"## The equations\n",
|
||
"\n",
|
||
"The Newton-Raphson formula consists geometrically of extending the\n",
|
||
"tangent line at a current point until it crosses zero, then setting\n",
|
||
"the next guess to the abscissa of that zero-crossing. The mathematics\n",
|
||
"behind this method is rather simple. Employing a Taylor expansion for\n",
|
||
"$x$ sufficiently close to the solution $s$, we have"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"<!-- Equation labels as ordinary links -->\n",
|
||
"<div id=\"eq:taylornr\"></div>\n",
|
||
"\n",
|
||
"$$\n",
|
||
"f(s)=0=f(x)+(s-x)f'(x)+\\frac{(s-x)^2}{2}f''(x) +\\dots.\n",
|
||
" \\label{eq:taylornr} \\tag{1}\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"For small enough values of the function and for well-behaved\n",
|
||
"functions, the terms beyond linear are unimportant, hence we obtain"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"f(x)+(s-x)f'(x)\\approx 0,\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"yielding"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"s\\approx x-\\frac{f(x)}{f'(x)}.\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"Having in mind an iterative procedure, it is natural to start iterating with"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"x_{n+1}=x_n-\\frac{f(x_n)}{f'(x_n)}.\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"## Simple geometric interpretation\n",
|
||
"\n",
|
||
"The above is Newton-Raphson's method. It has a simple geometric\n",
|
||
"interpretation, namely $x_{n+1}$ is the point where the tangent from\n",
|
||
"$(x_n,f(x_n))$ crosses the $x-$axis. Close to the solution,\n",
|
||
"Newton-Raphson converges fast to the desired result. However, if we\n",
|
||
"are far from a root, where the higher-order terms in the series are\n",
|
||
"important, the Newton-Raphson formula can give grossly inaccurate\n",
|
||
"results. For instance, the initial guess for the root might be so far\n",
|
||
"from the true root as to let the search interval include a local\n",
|
||
"maximum or minimum of the function. If an iteration places a trial\n",
|
||
"guess near such a local extremum, so that the first derivative nearly\n",
|
||
"vanishes, then Newton-Raphson may fail totally\n",
|
||
"\n",
|
||
"\n",
|
||
"## Extending to more than one variable\n",
|
||
"\n",
|
||
"Newton's method can be generalized to systems of several non-linear equations\n",
|
||
"and variables. Consider the case with two equations"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\begin{array}{cc} f_1(x_1,x_2) &=0\\\\\n",
|
||
" f_2(x_1,x_2) &=0\\end{array},\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"which we Taylor expand to obtain"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\begin{array}{cc} 0=f_1(x_1+h_1,x_2+h_2)=&f_1(x_1,x_2)+h_1\n",
|
||
" \\partial f_1/\\partial x_1+h_2\n",
|
||
" \\partial f_1/\\partial x_2+\\dots\\\\\n",
|
||
" 0=f_2(x_1+h_1,x_2+h_2)=&f_2(x_1,x_2)+h_1\n",
|
||
" \\partial f_2/\\partial x_1+h_2\n",
|
||
" \\partial f_2/\\partial x_2+\\dots\n",
|
||
" \\end{array}.\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"Defining the Jacobian matrix ${\\bf \\hat{J}}$ we have"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"{\\bf \\hat{J}}=\\left( \\begin{array}{cc}\n",
|
||
" \\partial f_1/\\partial x_1 & \\partial f_1/\\partial x_2 \\\\\n",
|
||
" \\partial f_2/\\partial x_1 &\\partial f_2/\\partial x_2\n",
|
||
" \\end{array} \\right),\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"we can rephrase Newton's method as"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\left(\\begin{array}{c} x_1^{n+1} \\\\ x_2^{n+1} \\end{array} \\right)=\n",
|
||
"\\left(\\begin{array}{c} x_1^{n} \\\\ x_2^{n} \\end{array} \\right)+\n",
|
||
"\\left(\\begin{array}{c} h_1^{n} \\\\ h_2^{n} \\end{array} \\right),\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"where we have defined"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\left(\\begin{array}{c} h_1^{n} \\\\ h_2^{n} \\end{array} \\right)=\n",
|
||
" -{\\bf \\hat{J}}^{-1}\n",
|
||
" \\left(\\begin{array}{c} f_1(x_1^{n},x_2^{n}) \\\\ f_2(x_1^{n},x_2^{n}) \\end{array} \\right).\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"We need thus to compute the inverse of the Jacobian matrix and it\n",
|
||
"is to understand that difficulties may\n",
|
||
"arise in case ${\\bf \\hat{J}}$ is nearly singular.\n",
|
||
"\n",
|
||
"It is rather straightforward to extend the above scheme to systems of\n",
|
||
"more than two non-linear equations. In our case, the Jacobian matrix is given by the Hessian that represents the second derivative of cost function. \n",
|
||
"\n",
|
||
"\n",
|
||
"\n",
|
||
"## Steepest descent\n",
|
||
"\n",
|
||
"The method of steepest descent The basic idea of gradient descent is\n",
|
||
"that a function $F(\\mathbf{x})$, \n",
|
||
"$\\mathbf{x} \\equiv (x_1,\\cdots,x_n)$, decreases fastest if one goes from $\\bf {x}$ in the\n",
|
||
"direction of the negative gradient $-\\nabla F(\\mathbf{x})$.\n",
|
||
"\n",
|
||
"It can be shown that if"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\mathbf{x}_{k+1} = \\mathbf{x}_k - \\gamma_k \\nabla F(\\mathbf{x}_k),\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"with $\\gamma_k > 0$.\n",
|
||
"\n",
|
||
"For $\\gamma_k$ small enough, then $F(\\mathbf{x}_{k+1}) \\leq\n",
|
||
"F(\\mathbf{x}_k)$. This means that for a sufficiently small $\\gamma_k$\n",
|
||
"we are always moving towards smaller function values, i.e a minimum.\n",
|
||
"\n",
|
||
"<!-- !split -->\n",
|
||
"## More on Steepest descent\n",
|
||
"\n",
|
||
"The previous observation is the basis of the method of steepest\n",
|
||
"descent, which is also referred to as just gradient descent (GD). One\n",
|
||
"starts with an initial guess $\\mathbf{x}_0$ for a minimum of $F$ and\n",
|
||
"computes new approximations according to"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\mathbf{x}_{k+1} = \\mathbf{x}_k - \\gamma_k \\nabla F(\\mathbf{x}_k), \\ \\ k \\geq 0.\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"The parameter $\\gamma_k$ is often referred to as the step length or\n",
|
||
"the learning rate within the context of Machine Learning.\n",
|
||
"\n",
|
||
"<!-- !split -->\n",
|
||
"## The ideal\n",
|
||
"\n",
|
||
"Ideally the sequence $\\{\\mathbf{x}_k \\}_{k=0}$ converges to a global\n",
|
||
"minimum of the function $F$. In general we do not know if we are in a\n",
|
||
"global or local minimum. In the special case when $F$ is a convex\n",
|
||
"function, all local minima are also global minima, so in this case\n",
|
||
"gradient descent can converge to the global solution. The advantage of\n",
|
||
"this scheme is that it is conceptually simple and straightforward to\n",
|
||
"implement. However the method in this form has some severe\n",
|
||
"limitations:\n",
|
||
"\n",
|
||
"In machine learing we are often faced with non-convex high dimensional\n",
|
||
"cost functions with many local minima. Since GD is deterministic we\n",
|
||
"will get stuck in a local minimum, if the method converges, unless we\n",
|
||
"have a very good intial guess. This also implies that the scheme is\n",
|
||
"sensitive to the chosen initial condition.\n",
|
||
"\n",
|
||
"Note that the gradient is a function of $\\mathbf{x} =\n",
|
||
"(x_1,\\cdots,x_n)$ which makes it expensive to compute numerically.\n",
|
||
"\n",
|
||
"\n",
|
||
"<!-- !split -->\n",
|
||
"## The sensitiveness of the gradient descent\n",
|
||
"\n",
|
||
"The gradient descent method \n",
|
||
"is sensitive to the choice of learning rate $\\gamma_k$. This is due\n",
|
||
"to the fact that we are only guaranteed that $F(\\mathbf{x}_{k+1}) \\leq\n",
|
||
"F(\\mathbf{x}_k)$ for sufficiently small $\\gamma_k$. The problem is to\n",
|
||
"determine an optimal learning rate. If the learning rate is chosen too\n",
|
||
"small the method will take a long time to converge and if it is too\n",
|
||
"large we can experience erratic behavior.\n",
|
||
"\n",
|
||
"Many of these shortcomings can be alleviated by introducing\n",
|
||
"randomness. One such method is that of Stochastic Gradient Descent\n",
|
||
"(SGD), see below.\n",
|
||
"\n",
|
||
"\n",
|
||
"<!-- !split -->\n",
|
||
"## Convex functions\n",
|
||
"\n",
|
||
"Ideally we want our cost/loss function to be convex(concave).\n",
|
||
"\n",
|
||
"First we give the definition of a convex set: A set $C$ in\n",
|
||
"$\\mathbb{R}^n$ is said to be convex if, for all $x$ and $y$ in $C$ and\n",
|
||
"all $t \\in (0,1)$ , the point $(1 − t)x + ty$ also belongs to\n",
|
||
"C. Geometrically this means that every point on the line segment\n",
|
||
"connecting $x$ and $y$ is in $C$ as discussed below.\n",
|
||
"\n",
|
||
"The convex subsets of $\\mathbb{R}$ are the intervals of\n",
|
||
"$\\mathbb{R}$. Examples of convex sets of $\\mathbb{R}^2$ are the\n",
|
||
"regular polygons (triangles, rectangles, pentagons, etc...).\n",
|
||
"\n",
|
||
"## Convex function\n",
|
||
"\n",
|
||
"**Convex function**: Let $X \\subset \\mathbb{R}^n$ be a convex set. Assume that the function $f: X \\rightarrow \\mathbb{R}$ is continuous, then $f$ is said to be convex if $$f(tx_1 + (1-t)x_2) \\leq tf(x_1) + (1-t)f(x_2) $$ for all $x_1, x_2 \\in X$ and for all $t \\in [0,1]$. If $\\leq$ is replaced with a strict inequaltiy in the definition, we demand $x_1 \\neq x_2$ and $t\\in(0,1)$ then $f$ is said to be strictly convex. For a single variable function, convexity means that if you draw a straight line connecting $f(x_1)$ and $f(x_2)$, the value of the function on the interval $[x_1,x_2]$ is always below the line as illustrated below.\n",
|
||
"\n",
|
||
"## Conditions on convex functions\n",
|
||
"\n",
|
||
"In the following we state first and second-order conditions which\n",
|
||
"ensures convexity of a function $f$. We write $D_f$ to denote the\n",
|
||
"domain of $f$, i.e the subset of $R^n$ where $f$ is defined. For more\n",
|
||
"details and proofs we refer to: [S. Boyd and L. Vandenberghe. Convex Optimization. Cambridge University Press](http://stanford.edu/boyd/cvxbook/, 2004).\n",
|
||
"\n",
|
||
"**First order condition.**\n",
|
||
"\n",
|
||
"Suppose $f$ is differentiable (i.e $\\nabla f(x)$ is well defined for\n",
|
||
"all $x$ in the domain of $f$). Then $f$ is convex if and only if $D_f$\n",
|
||
"is a convex set and $$f(y) \\geq f(x) + \\nabla f(x)^T (y-x) $$ holds\n",
|
||
"for all $x,y \\in D_f$. This condition means that for a convex function\n",
|
||
"the first order Taylor expansion (right hand side above) at any point\n",
|
||
"a global under estimator of the function. To convince yourself you can\n",
|
||
"make a drawing of $f(x) = x^2+1$ and draw the tangent line to $f(x)$ and\n",
|
||
"note that it is always below the graph.\n",
|
||
"\n",
|
||
"\n",
|
||
"\n",
|
||
"**Second order condition.**\n",
|
||
"\n",
|
||
"Assume that $f$ is twice\n",
|
||
"differentiable, i.e the Hessian matrix exists at each point in\n",
|
||
"$D_f$. Then $f$ is convex if and only if $D_f$ is a convex set and its\n",
|
||
"Hessian is positive semi-definite for all $x\\in D_f$. For a\n",
|
||
"single-variable function this reduces to $f''(x) \\geq 0$. Geometrically this means that $f$ has nonnegative curvature\n",
|
||
"everywhere.\n",
|
||
"\n",
|
||
"\n",
|
||
"\n",
|
||
"This condition is particularly useful since it gives us an procedure for determining if the function under consideration is convex, apart from using the definition.\n",
|
||
"\n",
|
||
"## More on convex functions\n",
|
||
"\n",
|
||
"The next result is of great importance to us and the reason why we are\n",
|
||
"going on about convex functions. In machine learning we frequently\n",
|
||
"have to minimize a loss/cost function in order to find the best\n",
|
||
"parameters for the model we are considering. \n",
|
||
"\n",
|
||
"Ideally we want the\n",
|
||
"global minimum (for high-dimensional models it is hard to know\n",
|
||
"if we have local or global minimum). However, if the cost/loss function\n",
|
||
"is convex the following result provides invaluable information:\n",
|
||
"\n",
|
||
"**Any minimum is global for convex functions.**\n",
|
||
"\n",
|
||
"Consider the problem of finding $x \\in \\mathbb{R}^n$ such that $f(x)$\n",
|
||
"is minimal, where $f$ is convex and differentiable. Then, any point\n",
|
||
"$x^*$ that satisfies $\\nabla f(x^*) = 0$ is a global minimum.\n",
|
||
"\n",
|
||
"\n",
|
||
"\n",
|
||
"This result means that if we know that the cost/loss function is convex and we are able to find a minimum, we are guaranteed that it is a global minimum.\n",
|
||
"\n",
|
||
"## Some simple problems\n",
|
||
"\n",
|
||
"1. Show that $f(x)=x^2$ is convex for $x \\in \\mathbb{R}$ using the definition of convexity. Hint: If you re-write the definition, $f$ is convex if the following holds for all $x,y \\in D_f$ and any $\\lambda \\in [0,1]$ $\\lambda f(x)+(1-\\lambda)f(y)-f(\\lambda x + (1-\\lambda) y ) \\geq 0$.\n",
|
||
"\n",
|
||
"2. Using the second order condition show that the following functions are convex on the specified domain.\n",
|
||
"\n",
|
||
" * $f(x) = e^x$ is convex for $x \\in \\mathbb{R}$.\n",
|
||
"\n",
|
||
" * $g(x) = -\\ln(x)$ is convex for $x \\in (0,\\infty)$.\n",
|
||
"\n",
|
||
"\n",
|
||
"3. Let $f(x) = x^2$ and $g(x) = e^x$. Show that $f(g(x))$ and $g(f(x))$ is convex for $x \\in \\mathbb{R}$. Also show that if $f(x)$ is any convex function than $h(x) = e^{f(x)}$ is convex.\n",
|
||
"\n",
|
||
"4. A norm is any function that satisfy the following properties\n",
|
||
"\n",
|
||
" * $f(\\alpha x) = |\\alpha| f(x)$ for all $\\alpha \\in \\mathbb{R}$.\n",
|
||
"\n",
|
||
" * $f(x+y) \\leq f(x) + f(y)$\n",
|
||
"\n",
|
||
" * $f(x) \\leq 0$ for all $x \\in \\mathbb{R}^n$ with equality if and only if $x = 0$\n",
|
||
"\n",
|
||
"\n",
|
||
"Using the definition of convexity, try to show that a function satisfying the properties above is convex (the third condition is not needed to show this).\n",
|
||
"\n",
|
||
"\n",
|
||
"## Standard steepest descent\n",
|
||
"\n",
|
||
"\n",
|
||
"Before we proceed, we would like to mention the approach called the **standard Steepest descent**, which again leads to us having to be able to compute a matrix.\n",
|
||
"\n",
|
||
"[The success of the CG method](https://www.cs.cmu.edu/~quake-papers/painless-conjugate-gradient.pdf)\n",
|
||
"for finding solutions of non-linear problems is based on the theory\n",
|
||
"of conjugate gradients for linear systems of equations. It belongs to\n",
|
||
"the class of iterative methods for solving problems from linear\n",
|
||
"algebra of the type"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{A}\\hat{x} = \\hat{b}.\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"In the iterative process we end up with a problem like"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{r}= \\hat{b}-\\hat{A}\\hat{x},\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"where $\\hat{r}$ is the so-called residual or error in the iterative process.\n",
|
||
"\n",
|
||
"When we have found the exact solution, $\\hat{r}=0$.\n",
|
||
"\n",
|
||
"## Gradient method\n",
|
||
"\n",
|
||
"The residual is zero when we reach the minimum of the quadratic equation"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"P(\\hat{x})=\\frac{1}{2}\\hat{x}^T\\hat{A}\\hat{x} - \\hat{x}^T\\hat{b},\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"with the constraint that the matrix $\\hat{A}$ is positive definite and\n",
|
||
"symmetric. If we search for a minimum of the quantum mechanical\n",
|
||
"variance, then the matrix $\\hat{A}$, which is called the Hessian, is\n",
|
||
"given by the second-derivative of the function we want to minimize.\n",
|
||
"This quantity is always positive definite. \n",
|
||
"\n",
|
||
"\n",
|
||
"## Steepest descent method\n",
|
||
"\n",
|
||
"We denote the initial guess for $\\hat{x}$ as $\\hat{x}_0$. \n",
|
||
"We can assume without loss of generality that"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{x}_0=0,\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"or consider the system"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{A}\\hat{z} = \\hat{b}-\\hat{A}\\hat{x}_0,\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"instead.\n",
|
||
"\n",
|
||
"\n",
|
||
"## Steepest descent method\n",
|
||
"One can show that the solution $\\hat{x}$ is also the unique minimizer of the quadratic form"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"f(\\hat{x}) = \\frac{1}{2}\\hat{x}^T\\hat{A}\\hat{x} - \\hat{x}^T \\hat{x} , \\quad \\hat{x}\\in\\mathbf{R}^n.\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"This suggests taking the first basis vector $\\hat{p}_1$ \n",
|
||
"to be the gradient of $f$ at $\\hat{x}=\\hat{x}_0$, \n",
|
||
"which equals"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{A}\\hat{x}_0-\\hat{b},\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"and \n",
|
||
"$\\hat{x}_0=0$ it is equal $-\\hat{b}$.\n",
|
||
"\n",
|
||
"\n",
|
||
"\n",
|
||
"\n",
|
||
"## Gradient descent method\n",
|
||
"Let $\\hat{r}_k$ be the residual at the $k$-th step:"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{r}_k=\\hat{b}-\\hat{A}\\hat{x}_k.\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"Note that $\\hat{r}_k$ is the negative gradient of $f$ at \n",
|
||
"$\\hat{x}=\\hat{x}_k$, \n",
|
||
"so the gradient descent method would be to move in the direction $\\hat{r}_k$. \n",
|
||
"This gives the following expression"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{p}_{k+1}=\\hat{r}_k-\\frac{\\hat{p}_k^T \\hat{A}\\hat{r}_k}{\\hat{p}_k^T\\hat{A}\\hat{p}_k} \\hat{p}_k.\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"## Final expressions\n",
|
||
"We can also compute the residual iteratively as"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{r}_{k+1}=\\hat{b}-\\hat{A}\\hat{x}_{k+1},\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"which equals"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{b}-\\hat{A}(\\hat{x}_k+\\alpha_k\\hat{p}_k),\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"or"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"(\\hat{b}-\\hat{A}\\hat{x}_k)-\\alpha_k\\hat{A}\\hat{p}_k,\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"which gives"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{r}_{k+1}=\\hat{r}_k-\\hat{A}\\hat{p}_{k},\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"## The Steepest descent algorithm\n",
|
||
"\n",
|
||
"\n",
|
||
"## Simple codes for steepest descent and conjugate gradient using a $2\\times 2$ matrix, in c++, Python code to come"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
" #include <cmath>\n",
|
||
" #include <iostream>\n",
|
||
" #include <fstream>\n",
|
||
" #include <iomanip>\n",
|
||
" #include \"vectormatrixclass.h\"\n",
|
||
" using namespace std;\n",
|
||
" // Main function begins here\n",
|
||
" int main(int argc, char * argv[]){\n",
|
||
" int dim = 2;\n",
|
||
" Vector x(dim),xsd(dim), b(dim),x0(dim);\n",
|
||
" Matrix A(dim,dim);\n",
|
||
" \n",
|
||
" // Set our initial guess\n",
|
||
" x0(0) = x0(1) = 0;\n",
|
||
" // Set the matrix\n",
|
||
" A(0,0) = 3; A(1,0) = 2; A(0,1) = 2; A(1,1) = 6;\n",
|
||
" b(0) = 2; b(1) = -8;\n",
|
||
" cout << \"The Matrix A that we are using: \" << endl;\n",
|
||
" A.Print();\n",
|
||
" cout << endl;\n",
|
||
" xsd = SteepestDescent(A,b,x0);\n",
|
||
" cout << \"The approximate solution using Steepest Descent is: \" << endl;\n",
|
||
" xsd.Print();\n",
|
||
" cout << endl;\n",
|
||
" }\n"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"## The routine for the steepest descent method"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
" Vector SteepestDescent(Matrix A, Vector b, Vector x0){\n",
|
||
" int IterMax, i;\n",
|
||
" int dim = x0.Dimension();\n",
|
||
" const double tolerance = 1.0e-14;\n",
|
||
" Vector x(dim),f(dim),z(dim);\n",
|
||
" double c,alpha,d;\n",
|
||
" IterMax = 30;\n",
|
||
" x = x0;\n",
|
||
" f = A*x-b;\n",
|
||
" i = 0;\n",
|
||
" while (i <= IterMax){\n",
|
||
" z = A*f;\n",
|
||
" c = dot(f,f);\n",
|
||
" alpha = c/dot(f,z);\n",
|
||
" x = x - alpha*f;\n",
|
||
" f = A*x-b;\n",
|
||
" if(sqrt(dot(f,f)) < tolerance) break;\n",
|
||
" i++;\n",
|
||
" }\n",
|
||
" return x;\n",
|
||
" }\n"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"<!-- !split -->\n",
|
||
"## Revisiting our first homework\n",
|
||
"\n",
|
||
"We will use linear regression as a case study for the gradient descent\n",
|
||
"methods. Linear regression is a great test case for the gradient\n",
|
||
"descent methods discussed in the lectures since it has several\n",
|
||
"desirable properties such as:\n",
|
||
"\n",
|
||
"1. An analytical solution (recall homework set 1).\n",
|
||
"\n",
|
||
"2. The gradient can be computed analytically.\n",
|
||
"\n",
|
||
"3. The cost function is convex which guarantees that gradient descent converges for small enough learning rates\n",
|
||
"\n",
|
||
"We revisit the example from homework set 1 where we had"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"y_i = 5x_i^2 + 0.1\\xi_i, \\ i=1,\\cdots,100\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"with $x_i \\in [0,1] $ chosen randomly with a uniform distribution. Additionally $\\xi_i$ represents stochastic noise chosen according to a normal distribution $\\cal {N}(0,1)$. \n",
|
||
"The linear regression model is given by"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"h_\\beta(x) = \\hat{y} = \\beta_0 + \\beta_1 x,\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"such that"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{y}_i = \\beta_0 + \\beta_1 x_i.\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"<!-- !split -->\n",
|
||
"## Gradient descent example\n",
|
||
"\n",
|
||
"Let $\\mathbf{y} = (y_1,\\cdots,y_n)^T$, $\\mathbf{\\hat{y}} = (\\hat{y}_1,\\cdots,\\hat{y}_n)^T$ and $\\beta = (\\beta_0, \\beta_1)^T$\n",
|
||
"\n",
|
||
"It is convenient to write $\\mathbf{\\hat{y}} = X\\beta$ where $X \\in \\mathbb{R}^{100 \\times 2} $ is the design matrix given by"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"X \\equiv \\begin{bmatrix}\n",
|
||
"1 & x_1 \\\\\n",
|
||
"\\vdots & \\vdots \\\\\n",
|
||
"1 & x_{100} & \\\\\n",
|
||
"\\end{bmatrix}.\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"The loss function is given by"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"C(\\beta) = ||X\\beta-\\mathbf{y}||^2 = ||X\\beta||^2 - 2 \\mathbf{y}^T X\\beta + ||\\mathbf{y}||^2 = \\sum_{i=1}^{100} (\\beta_0 + \\beta_1 x_i)^2 - 2 y_i (\\beta_0 + \\beta_1 x_i) + y_i^2\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"and we want to find $\\beta$ such that $C(\\beta)$ is minimized.\n",
|
||
"\n",
|
||
"## The derivative of the cost/loss function\n",
|
||
"\n",
|
||
"Computing $\\partial C(\\beta) / \\partial \\beta_0$ and $\\partial C(\\beta) / \\partial \\beta_1$ we can show that the gradient can be written as"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\nabla_{\\beta} C(\\beta) = (\\partial C(\\beta) / \\partial \\beta_0, \\partial C(\\beta) / \\partial \\beta_1)^T = 2\\begin{bmatrix} \\sum_{i=1}^{100} \\left(\\beta_0+\\beta_1x_i-y_i\\right) \\\\\n",
|
||
"\\sum_{i=1}^{100}\\left( x_i (\\beta_0+\\beta_1x_i)-y_ix_i\\right) \\\\\n",
|
||
"\\end{bmatrix} = 2X^T(X\\beta - \\mathbf{y}),\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"where $X$ is the design matrix defined above.\n",
|
||
"\n",
|
||
"## The Hessian matrix\n",
|
||
"The Hessian matrix of $C(\\beta)$ is given by"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{H} \\equiv \\begin{bmatrix}\n",
|
||
"\\frac{\\partial^2 C(\\beta)}{\\partial \\beta_0^2} & \\frac{\\partial^2 C(\\beta)}{\\partial \\beta_0 \\partial \\beta_1} \\\\\n",
|
||
"\\frac{\\partial^2 C(\\beta)}{\\partial \\beta_0 \\partial \\beta_1} & \\frac{\\partial^2 C(\\beta)}{\\partial \\beta_1^2} & \\\\\n",
|
||
"\\end{bmatrix} = 2X^T X.\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"This result implies that $C(\\beta)$ is a convex function since the matrix $X^T X$ always is positive semi-definite.\n",
|
||
"\n",
|
||
"## Simple program\n",
|
||
"\n",
|
||
"We can now write a program that minimizes $C(\\beta)$ using the gradient descent method with a constant learning rate $\\gamma$ according to"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\beta_{k+1} = \\beta_k - \\gamma \\nabla_\\beta C(\\beta_k), \\ k=0,1,\\cdots\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"We can use the expression we computed for the gradient and let use a\n",
|
||
"$\\beta_0$ be chosen randomly and let $\\gamma = 0.001$. Stop iterating\n",
|
||
"when $||\\nabla_\\beta C(\\beta_k) || \\leq \\epsilon = 10^{-8}$. \n",
|
||
"\n",
|
||
"And finally we can compare our solution for $\\beta$ with the analytic result given by \n",
|
||
"$\\beta= (X^TX)^{-1} X^T \\mathbf{y}$."
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "code",
|
||
"execution_count": 1,
|
||
"metadata": {
|
||
"collapsed": false
|
||
},
|
||
"outputs": [],
|
||
"source": [
|
||
"import numpy as np\n",
|
||
"\n",
|
||
"\"\"\"\n",
|
||
"The following setup is just a suggestion, feel free to write it the way you like.\n",
|
||
"\"\"\"\n",
|
||
"\n",
|
||
"#Setup problem described in the exercise\n",
|
||
"N = 100 #Nr of datapoints\n",
|
||
"M = 2 #Nr of features\n",
|
||
"x = np.random.rand(N) #Uniformly generated x-values in [0,1]\n",
|
||
"y = 5*x**2 + 0.1*np.random.randn(N)\n",
|
||
"X = np.c_[np.ones(N),x] #Construct design matrix\n",
|
||
"\n",
|
||
"#Compute beta according to normal equations to compare with GD solution\n",
|
||
"Xt_X_inv = np.linalg.inv(np.dot(X.T,X))\n",
|
||
"Xt_y = np.dot(X.transpose(),y)\n",
|
||
"beta_NE = np.dot(Xt_X_inv,Xt_y)\n",
|
||
"print(beta_NE)"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"## Gradient Descent Example\n",
|
||
"\n",
|
||
"Another simple example is here"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "code",
|
||
"execution_count": 2,
|
||
"metadata": {
|
||
"collapsed": false
|
||
},
|
||
"outputs": [],
|
||
"source": [
|
||
"%matplotlib inline\n",
|
||
"\n",
|
||
"\n",
|
||
"# Importing various packages\n",
|
||
"from random import random, seed\n",
|
||
"import numpy as np\n",
|
||
"import matplotlib.pyplot as plt\n",
|
||
"from mpl_toolkits.mplot3d import Axes3D\n",
|
||
"from matplotlib import cm\n",
|
||
"from matplotlib.ticker import LinearLocator, FormatStrFormatter\n",
|
||
"import sys\n",
|
||
"\n",
|
||
"x = 2*np.random.rand(100,1)\n",
|
||
"y = 4+3*x+np.random.randn(100,1)\n",
|
||
"\n",
|
||
"xb = np.c_[np.ones((100,1)), x]\n",
|
||
"beta_linreg = np.linalg.inv(xb.T.dot(xb)).dot(xb.T).dot(y)\n",
|
||
"print(beta_linreg)\n",
|
||
"beta = np.random.randn(2,1)\n",
|
||
"\n",
|
||
"eta = 0.1\n",
|
||
"Niterations = 1000\n",
|
||
"m = 100\n",
|
||
"\n",
|
||
"for iter in range(Niterations):\n",
|
||
" gradients = 2.0/m*xb.T.dot(xb.dot(beta)-y)\n",
|
||
" beta -= eta*gradients\n",
|
||
"\n",
|
||
"print(beta)\n",
|
||
"xnew = np.array([[0],[2]])\n",
|
||
"xbnew = np.c_[np.ones((2,1)), xnew]\n",
|
||
"ypredict = xbnew.dot(beta)\n",
|
||
"ypredict2 = xbnew.dot(beta_linreg)\n",
|
||
"plt.plot(xnew, ypredict, \"r-\")\n",
|
||
"plt.plot(xnew, ypredict2, \"b-\")\n",
|
||
"plt.plot(x, y ,'ro')\n",
|
||
"plt.axis([0,2.0,0, 15.0])\n",
|
||
"plt.xlabel(r'$x$')\n",
|
||
"plt.ylabel(r'$y$')\n",
|
||
"plt.title(r'Gradient descent example')\n",
|
||
"plt.show()"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"## And a corresponding example using **scikit-learn**"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "code",
|
||
"execution_count": 3,
|
||
"metadata": {
|
||
"collapsed": false
|
||
},
|
||
"outputs": [],
|
||
"source": [
|
||
"# Importing various packages\n",
|
||
"from random import random, seed\n",
|
||
"import numpy as np\n",
|
||
"import matplotlib.pyplot as plt\n",
|
||
"from sklearn.linear_model import SGDRegressor\n",
|
||
"\n",
|
||
"x = 2*np.random.rand(100,1)\n",
|
||
"y = 4+3*x+np.random.randn(100,1)\n",
|
||
"\n",
|
||
"xb = np.c_[np.ones((100,1)), x]\n",
|
||
"beta_linreg = np.linalg.inv(xb.T.dot(xb)).dot(xb.T).dot(y)\n",
|
||
"print(beta_linreg)\n",
|
||
"sgdreg = SGDRegressor(n_iter = 50, penalty=None, eta0=0.1)\n",
|
||
"sgdreg.fit(x,y.ravel())\n",
|
||
"print(sgdreg.intercept_, sgdreg.coef_)"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"<!-- !split -->\n",
|
||
"## Gradient descent and Ridge\n",
|
||
"\n",
|
||
"We have also discussed Ridge regression where the loss function contains a regularized given by the $L_2$ norm of $\\beta$,"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"C_{\\text{ridge}}(\\beta) = ||X\\beta -\\mathbf{y}||^2 + \\lambda ||\\beta||^2, \\ \\lambda \\geq 0.\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"In order to minimize $C_{\\text{ridge}}(\\beta)$ using GD we only have adjust the gradient as follows"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\nabla_\\beta C_{\\text{ridge}}(\\beta) = 2\\begin{bmatrix} \\sum_{i=1}^{100} \\left(\\beta_0+\\beta_1x_i-y_i\\right) \\\\\n",
|
||
"\\sum_{i=1}^{100}\\left( x_i (\\beta_0+\\beta_1x_i)-y_ix_i\\right) \\\\\n",
|
||
"\\end{bmatrix} + 2\\lambda\\begin{bmatrix} \\beta_0 \\\\ \\beta_1\\end{bmatrix} = 2 (X^T(X\\beta - \\mathbf{y})+\\lambda \\beta).\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"We can now extend our program to minimize $C_{\\text{ridge}}(\\beta)$ using gradient descent and compare with the analytical solution given by"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\beta_{\\text{ridge}} = \\left(X^T X + \\lambda I_{2 \\times 2} \\right)^{-1} X^T \\mathbf{y},\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"for $\\lambda = {0,1,10,50,100}$ ($\\lambda = 0$ corresponds to ordinary least squares). \n",
|
||
"We can then compute $||\\beta_{\\text{ridge}}||$ for each $\\lambda$."
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "code",
|
||
"execution_count": 4,
|
||
"metadata": {
|
||
"collapsed": false
|
||
},
|
||
"outputs": [],
|
||
"source": [
|
||
"import numpy as np\n",
|
||
"\n",
|
||
"\"\"\"\n",
|
||
"The following setup is just a suggestion, feel free to write it the way you like.\n",
|
||
"\"\"\"\n",
|
||
"\n",
|
||
"#Setup problem described in the exercise\n",
|
||
"N = 100 #Nr of datapoints\n",
|
||
"M = 2 #Nr of features\n",
|
||
"x = np.random.rand(N)\n",
|
||
"y = 5*x**2 + 0.1*np.random.randn(N)\n",
|
||
"\n",
|
||
"\n",
|
||
"#Compute analytic beta for Ridge regression \n",
|
||
"X = np.c_[np.ones(N),x]\n",
|
||
"XT_X = np.dot(X.T,X)\n",
|
||
"\n",
|
||
"l = 0.1 #Ridge parameter lambda\n",
|
||
"Id = np.eye(XT_X.shape[0])\n",
|
||
"\n",
|
||
"Z = np.linalg.inv(XT_X+l*Id)\n",
|
||
"beta_ridge = np.dot(Z,np.dot(X.T,y))\n",
|
||
"\n",
|
||
"print(beta_ridge)\n",
|
||
"print(np.linalg.norm(beta_ridge)) #||beta||"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"## Stochastic Gradient Descent\n",
|
||
"\n",
|
||
"Stochastic gradient descent (SGD) and variants thereof address some of\n",
|
||
"the shortcomings of the Gradient descent method discussed above.\n",
|
||
"\n",
|
||
"The underlying idea of SGD comes from the observation that the cost\n",
|
||
"function, which we want to minimize, can almost always be written as a\n",
|
||
"sum over $n$ data points $\\{\\mathbf{x}_i\\}_{i=1}^n$,"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"C(\\mathbf{\\beta}) = \\sum_{i=1}^n c_i(\\mathbf{x}_i,\n",
|
||
"\\mathbf{\\beta}).\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"## Computation of gradients\n",
|
||
"\n",
|
||
"This in turn means that the gradient can be\n",
|
||
"computed as a sum over $i$-gradients"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\nabla_\\beta C(\\mathbf{\\beta}) = \\sum_i^n \\nabla_\\beta c_i(\\mathbf{x}_i,\n",
|
||
"\\mathbf{\\beta}).\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"Stochasticity/randomness is introduced by only taking the\n",
|
||
"gradient on a subset of the data called minibatches. If there are $n$\n",
|
||
"data points and the size of each minibatch is $M$, there will be $n/M$\n",
|
||
"minibatches. We denote these minibatches by $B_k$ where\n",
|
||
"$k=1,\\cdots,n/M$.\n",
|
||
"\n",
|
||
"## SGD example\n",
|
||
"As an example, suppose we have $10$ data points $(\\mathbf{x}_1,\\cdots, \\mathbf{x}_{10})$ \n",
|
||
"and we choose to have $M=5$ minibathces,\n",
|
||
"then each minibatch contains two data points. In particular we have\n",
|
||
"$B_1 = (\\mathbf{x}_1,\\mathbf{x}_2), \\cdots, B_5 =\n",
|
||
"(\\mathbf{x}_9,\\mathbf{x}_{10})$. Note that if you choose $M=1$ you\n",
|
||
"have only a single batch with all data points and on the other extreme,\n",
|
||
"you may choose $M=n$ resulting in a minibatch for each datapoint, i.e\n",
|
||
"$B_k = \\mathbf{x}_k$.\n",
|
||
"\n",
|
||
"The idea is now to approximate the gradient by replacing the sum over\n",
|
||
"all data points with a sum over the data points in one the minibatches\n",
|
||
"picked at random in each gradient descent step"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\nabla_{\\beta}\n",
|
||
"C(\\mathbf{\\beta}) = \\sum_{i=1}^n \\nabla_\\beta c_i(\\mathbf{x}_i,\n",
|
||
"\\mathbf{\\beta}) \\rightarrow \\sum_{i \\in B_k}^n \\nabla_\\beta\n",
|
||
"c_i(\\mathbf{x}_i, \\mathbf{\\beta}).\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"## The gradient step\n",
|
||
"\n",
|
||
"Thus a gradient descent step now looks like"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\beta_{j+1} = \\beta_j - \\gamma_j \\sum_{i \\in B_k}^n \\nabla_\\beta c_i(\\mathbf{x}_i,\n",
|
||
"\\mathbf{\\beta})\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"where $k$ is picked at random with equal\n",
|
||
"probability from $[1,n/M]$. An iteration over the number of\n",
|
||
"minibathces (n/M) is commonly referred to as an epoch. Thus it is\n",
|
||
"typical to choose a number of epochs and for each epoch iterate over\n",
|
||
"the number of minibatches, as exemplified in the code below.\n",
|
||
"\n",
|
||
"## Simple example code"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "code",
|
||
"execution_count": 5,
|
||
"metadata": {
|
||
"collapsed": false
|
||
},
|
||
"outputs": [],
|
||
"source": [
|
||
"import numpy as np \n",
|
||
"\n",
|
||
"n = 100 #100 datapoints \n",
|
||
"M = 5 #size of each minibatch\n",
|
||
"m = int(n/M) #number of minibatches\n",
|
||
"n_epochs = 10 #number of epochs\n",
|
||
"\n",
|
||
"j = 0\n",
|
||
"for epoch in range(1,n_epochs+1):\n",
|
||
" for i in range(m):\n",
|
||
" k = np.random.randint(m) #Pick the k-th minibatch at random\n",
|
||
" #Compute the gradient using the data in minibatch Bk\n",
|
||
" #Compute new suggestion for \n",
|
||
" j += 1"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"Taking the gradient only on a subset of the data has two important\n",
|
||
"benefits. First, it introduces randomness which decreases the chance\n",
|
||
"that our opmization scheme gets stuck in a local minima. Second, if\n",
|
||
"the size of the minibatches are small relative to the number of\n",
|
||
"datapoints ($M < n$), the computation of the gradient is much\n",
|
||
"cheaper since we sum over the datapoints in the $k-th$ minibatch and not\n",
|
||
"all $n$ datapoints.\n",
|
||
"\n",
|
||
"## When do we stop?\n",
|
||
"\n",
|
||
"A natural question is when do we stop the search for a new minimum?\n",
|
||
"One possibility is to compute the full gradient after a given number\n",
|
||
"of epochs and check if the norm of the gradient is smaller than some\n",
|
||
"threshold and stop if true. However, the condition that the gradient\n",
|
||
"is zero is valid also for local minima, so this would only tell us\n",
|
||
"that we are close to a local/global minimum. However, we could also\n",
|
||
"evaluate the cost function at this point, store the result and\n",
|
||
"continue the search. If the test kicks in at a later stage we can\n",
|
||
"compare the values of the cost function and keep the $\\beta$ that\n",
|
||
"gave the lowest value.\n",
|
||
"\n",
|
||
"## Slightly different approach\n",
|
||
"\n",
|
||
"Another approach is to let the step length $\\gamma_j$ depend on the\n",
|
||
"number of epochs in such a way that it becomes very small after a\n",
|
||
"reasonable time such that we do not move at all.\n",
|
||
"\n",
|
||
"As an example, let $e = 0,1,2,3,\\cdots$ denote the current epoch and let $t_0, t_1 > 0$ be two fixed numbers. Furthermore, let $t = e \\cdot m + i$ where $m$ is the number of minibatches and $i=0,\\cdots,m-1$. Then the function $$\\gamma_j(t; t_0, t_1) = \\frac{t_0}{t+t_1} $$ goes to zero as the number of epochs gets large. I.e. we start with a step length $\\gamma_j (0; t_0, t_1) = t_0/t_1$ which decays in *time* $t$.\n",
|
||
"\n",
|
||
"In this way we can fix the number of epochs, compute $\\beta$ and\n",
|
||
"evaluate the cost function at the end. Repeating the computation will\n",
|
||
"give a different result since the scheme is random by design. Then we\n",
|
||
"pick the final $\\beta$ that gives the lowest value of the cost\n",
|
||
"function."
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "code",
|
||
"execution_count": 6,
|
||
"metadata": {
|
||
"collapsed": false
|
||
},
|
||
"outputs": [],
|
||
"source": [
|
||
"import numpy as np \n",
|
||
"\n",
|
||
"def step_length(t,t0,t1):\n",
|
||
" return t0/(t+t1)\n",
|
||
"\n",
|
||
"n = 100 #100 datapoints \n",
|
||
"M = 5 #size of each minibatch\n",
|
||
"m = int(n/M) #number of minibatches\n",
|
||
"n_epochs = 500 #number of epochs\n",
|
||
"t0 = 1.0\n",
|
||
"t1 = 10\n",
|
||
"\n",
|
||
"gamma_j = t0/t1\n",
|
||
"j = 0\n",
|
||
"for epoch in range(1,n_epochs+1):\n",
|
||
" for i in range(m):\n",
|
||
" k = np.random.randint(m) #Pick the k-th minibatch at random\n",
|
||
" #Compute the gradient using the data in minibatch Bk\n",
|
||
" #Compute new suggestion for beta\n",
|
||
" t = epoch*m+i\n",
|
||
" gamma_j = step_length(t,t0,t1)\n",
|
||
" j += 1\n",
|
||
"\n",
|
||
"print(\"gamma_j after %d epochs: %g\" % (n_epochs,gamma_j))"
|
||
]
|
||
}
|
||
],
|
||
"metadata": {},
|
||
"nbformat": 4,
|
||
"nbformat_minor": 2
|
||
}
|