2606 lines
74 KiB
Plaintext
2606 lines
74 KiB
Plaintext
{
|
||
"cells": [
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"<!-- dom:TITLE: Data Analysis and Machine Learning Lectures: Optimization and Gradient Methods -->\n",
|
||
"# Data Analysis and Machine Learning Lectures: Optimization and Gradient Methods\n",
|
||
"<!-- dom:AUTHOR: Morten Hjorth-Jensen at Department of Physics, University of Oslo & Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University -->\n",
|
||
"<!-- Author: --> \n",
|
||
"**Morten Hjorth-Jensen**, Department of Physics, University of Oslo and Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University\n",
|
||
"\n",
|
||
"Date: **Oct 6, 2018**\n",
|
||
"\n",
|
||
"Copyright 1999-2018, Morten Hjorth-Jensen. Released under CC Attribution-NonCommercial 4.0 license\n",
|
||
"\n",
|
||
"\n",
|
||
"\n",
|
||
"\n",
|
||
"## Optimization, the central part of any Machine Learning algortithm\n",
|
||
"\n",
|
||
"Almost every problem in machine learning and data science starts with\n",
|
||
"a dataset $X$, a model $g(\\beta)$, which is a function of the\n",
|
||
"parameters $\\beta$ and a cost function $C(X, g(\\beta))$ that allows\n",
|
||
"us to judge how well the model $g(\\beta)$ explains the observations\n",
|
||
"$X$. The model is fit by finding the values of $\\beta$ that minimize\n",
|
||
"the cost function. Ideally we would be able to solve for $\\beta$\n",
|
||
"analytically, however this is not possible in general and we must use\n",
|
||
"some approximative/numerical method to compute the minimum.\n",
|
||
"\n",
|
||
"\n",
|
||
"## Revisiting our Logistic Regression case\n",
|
||
"\n",
|
||
"In our discussion on Logistic Regression we studied the \n",
|
||
"case of\n",
|
||
"two classes, with $y_i$ either\n",
|
||
"$0$ or $1$. Furthermore we assumed also that we have only two\n",
|
||
"parameters $\\beta$ in our fitting, that is we\n",
|
||
"defined probabilities"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\begin{align*}\n",
|
||
"p(y_i=1|x_i,\\hat{\\beta}) &= \\frac{\\exp{(\\beta_0+\\beta_1x_i)}}{1+\\exp{(\\beta_0+\\beta_1x_i)}},\\nonumber\\\\\n",
|
||
"p(y_i=0|x_i,\\hat{\\beta}) &= 1 - p(y_i=1|x_i,\\hat{\\beta}),\n",
|
||
"\\end{align*}\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"where $\\hat{\\beta}$ are the weights we wish to extract from data, in our case $\\beta_0$ and $\\beta_1$. \n",
|
||
"\n",
|
||
"## The equations to solve\n",
|
||
"\n",
|
||
"Our compact equations used a definition of a vector $\\hat{y}$ with $n$\n",
|
||
"elements $y_i$, an $n\\times p$ matrix $\\hat{X}$ which contains the\n",
|
||
"$x_i$ values and a vector $\\hat{p}$ of fitted probabilities\n",
|
||
"$p(y_i\\vert x_i,\\hat{\\beta})$. We rewrote in a more compact form\n",
|
||
"the first derivative of the cost function as"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\frac{\\partial \\mathcal{C}(\\hat{\\beta})}{\\partial \\hat{\\beta}} = -\\hat{X}^T\\left(\\hat{y}-\\hat{p}\\right).\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"If we in addition define a diagonal matrix $\\hat{W}$ with elements \n",
|
||
"$p(y_i\\vert x_i,\\hat{\\beta})(1-p(y_i\\vert x_i,\\hat{\\beta})$, we can obtain a compact expression of the second derivative as"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\frac{\\partial^2 \\mathcal{C}(\\hat{\\beta})}{\\partial \\hat{\\beta}\\partial \\hat{\\beta}^T} = \\hat{X}^T\\hat{W}\\hat{X}.\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"This defines what is called the Hessian matrix.\n",
|
||
"\n",
|
||
"## Solving using Newton-Raphson's method\n",
|
||
"\n",
|
||
"If we can set up these equations, Newton-Raphson's iterative method is normally the method of choice. It requires however that we can compute in an efficient way the matrices that define the first and second derivatives. \n",
|
||
"\n",
|
||
"Our iterative scheme is then given by"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{\\beta}^{\\mathrm{new}} = \\hat{\\beta}^{\\mathrm{old}}-\\left(\\frac{\\partial^2 \\mathcal{C}(\\hat{\\beta})}{\\partial \\hat{\\beta}\\partial \\hat{\\beta}^T}\\right)^{-1}_{\\hat{\\beta}^{\\mathrm{old}}}\\times \\left(\\frac{\\partial \\mathcal{C}(\\hat{\\beta})}{\\partial \\hat{\\beta}}\\right)_{\\hat{\\beta}^{\\mathrm{old}}},\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"or in matrix form as"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{\\beta}^{\\mathrm{new}} = \\hat{\\beta}^{\\mathrm{old}}-\\left(\\hat{X}^T\\hat{W}\\hat{X} \\right)^{-1}\\times \\left(-\\hat{X}^T(\\hat{y}-\\hat{p}) \\right)_{\\hat{\\beta}^{\\mathrm{old}}}.\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"The right-hand side is computed with the old values of $\\beta$. \n",
|
||
"\n",
|
||
"If we can compute these matrices, in particular the Hessian, the above is often the easiest method to implement. \n",
|
||
"\n",
|
||
"\n",
|
||
"## Brief reminder on Newton-Raphson's method\n",
|
||
"\n",
|
||
"Let us quickly remind ourselves how we derive the above method.\n",
|
||
"\n",
|
||
"Perhaps the most celebrated of all one-dimensional root-finding\n",
|
||
"routines is Newton's method, also called the Newton-Raphson\n",
|
||
"method. This method requires the evaluation of both the\n",
|
||
"function $f$ and its derivative $f'$ at arbitrary points. \n",
|
||
"If you can only calculate the derivative\n",
|
||
"numerically and/or your function is not of the smooth type, we\n",
|
||
"normally discourage the use of this method.\n",
|
||
"\n",
|
||
"## The equations\n",
|
||
"\n",
|
||
"The Newton-Raphson formula consists geometrically of extending the\n",
|
||
"tangent line at a current point until it crosses zero, then setting\n",
|
||
"the next guess to the abscissa of that zero-crossing. The mathematics\n",
|
||
"behind this method is rather simple. Employing a Taylor expansion for\n",
|
||
"$x$ sufficiently close to the solution $s$, we have"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"<!-- Equation labels as ordinary links -->\n",
|
||
"<div id=\"eq:taylornr\"></div>\n",
|
||
"\n",
|
||
"$$\n",
|
||
"f(s)=0=f(x)+(s-x)f'(x)+\\frac{(s-x)^2}{2}f''(x) +\\dots.\n",
|
||
" \\label{eq:taylornr} \\tag{1}\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"For small enough values of the function and for well-behaved\n",
|
||
"functions, the terms beyond linear are unimportant, hence we obtain"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"f(x)+(s-x)f'(x)\\approx 0,\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"yielding"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"s\\approx x-\\frac{f(x)}{f'(x)}.\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"Having in mind an iterative procedure, it is natural to start iterating with"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"x_{n+1}=x_n-\\frac{f(x_n)}{f'(x_n)}.\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"## Simple geometric interpretation\n",
|
||
"\n",
|
||
"The above is Newton-Raphson's method. It has a simple geometric\n",
|
||
"interpretation, namely $x_{n+1}$ is the point where the tangent from\n",
|
||
"$(x_n,f(x_n))$ crosses the $x$-axis. Close to the solution,\n",
|
||
"Newton-Raphson converges fast to the desired result. However, if we\n",
|
||
"are far from a root, where the higher-order terms in the series are\n",
|
||
"important, the Newton-Raphson formula can give grossly inaccurate\n",
|
||
"results. For instance, the initial guess for the root might be so far\n",
|
||
"from the true root as to let the search interval include a local\n",
|
||
"maximum or minimum of the function. If an iteration places a trial\n",
|
||
"guess near such a local extremum, so that the first derivative nearly\n",
|
||
"vanishes, then Newton-Raphson may fail totally\n",
|
||
"\n",
|
||
"\n",
|
||
"## Extending to more than one variable\n",
|
||
"\n",
|
||
"Newton's method can be generalized to systems of several non-linear equations\n",
|
||
"and variables. Consider the case with two equations"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\begin{array}{cc} f_1(x_1,x_2) &=0\\\\\n",
|
||
" f_2(x_1,x_2) &=0,\\end{array}\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"which we Taylor expand to obtain"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\begin{array}{cc} 0=f_1(x_1+h_1,x_2+h_2)=&f_1(x_1,x_2)+h_1\n",
|
||
" \\partial f_1/\\partial x_1+h_2\n",
|
||
" \\partial f_1/\\partial x_2+\\dots\\\\\n",
|
||
" 0=f_2(x_1+h_1,x_2+h_2)=&f_2(x_1,x_2)+h_1\n",
|
||
" \\partial f_2/\\partial x_1+h_2\n",
|
||
" \\partial f_2/\\partial x_2+\\dots\n",
|
||
" \\end{array}.\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"Defining the Jacobian matrix ${\\bf \\hat{J}}$ we have"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"{\\bf \\hat{J}}=\\left( \\begin{array}{cc}\n",
|
||
" \\partial f_1/\\partial x_1 & \\partial f_1/\\partial x_2 \\\\\n",
|
||
" \\partial f_2/\\partial x_1 &\\partial f_2/\\partial x_2\n",
|
||
" \\end{array} \\right),\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"we can rephrase Newton's method as"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\left(\\begin{array}{c} x_1^{n+1} \\\\ x_2^{n+1} \\end{array} \\right)=\n",
|
||
"\\left(\\begin{array}{c} x_1^{n} \\\\ x_2^{n} \\end{array} \\right)+\n",
|
||
"\\left(\\begin{array}{c} h_1^{n} \\\\ h_2^{n} \\end{array} \\right),\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"where we have defined"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\left(\\begin{array}{c} h_1^{n} \\\\ h_2^{n} \\end{array} \\right)=\n",
|
||
" -{\\bf \\hat{J}}^{-1}\n",
|
||
" \\left(\\begin{array}{c} f_1(x_1^{n},x_2^{n}) \\\\ f_2(x_1^{n},x_2^{n}) \\end{array} \\right).\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"We need thus to compute the inverse of the Jacobian matrix and it\n",
|
||
"is to understand that difficulties may\n",
|
||
"arise in case ${\\bf \\hat{J}}$ is nearly singular.\n",
|
||
"\n",
|
||
"It is rather straightforward to extend the above scheme to systems of\n",
|
||
"more than two non-linear equations. In our case, the Jacobian matrix is given by the Hessian that represents the second derivative of cost function. \n",
|
||
"\n",
|
||
"\n",
|
||
"\n",
|
||
"## Steepest descent\n",
|
||
"\n",
|
||
"The basic idea of gradient descent is\n",
|
||
"that a function $F(\\mathbf{x})$, \n",
|
||
"$\\mathbf{x} \\equiv (x_1,\\cdots,x_n)$, decreases fastest if one goes from $\\bf {x}$ in the\n",
|
||
"direction of the negative gradient $-\\nabla F(\\mathbf{x})$.\n",
|
||
"\n",
|
||
"It can be shown that if"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\mathbf{x}_{k+1} = \\mathbf{x}_k - \\gamma_k \\nabla F(\\mathbf{x}_k),\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"with $\\gamma_k > 0$.\n",
|
||
"\n",
|
||
"For $\\gamma_k$ small enough, then $F(\\mathbf{x}_{k+1}) \\leq\n",
|
||
"F(\\mathbf{x}_k)$. This means that for a sufficiently small $\\gamma_k$\n",
|
||
"we are always moving towards smaller function values, i.e a minimum.\n",
|
||
"\n",
|
||
"<!-- !split -->\n",
|
||
"## More on Steepest descent\n",
|
||
"\n",
|
||
"The previous observation is the basis of the method of steepest\n",
|
||
"descent, which is also referred to as just gradient descent (GD). One\n",
|
||
"starts with an initial guess $\\mathbf{x}_0$ for a minimum of $F$ and\n",
|
||
"computes new approximations according to"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\mathbf{x}_{k+1} = \\mathbf{x}_k - \\gamma_k \\nabla F(\\mathbf{x}_k), \\ \\ k \\geq 0.\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"The parameter $\\gamma_k$ is often referred to as the step length or\n",
|
||
"the learning rate within the context of Machine Learning.\n",
|
||
"\n",
|
||
"<!-- !split -->\n",
|
||
"## The ideal\n",
|
||
"\n",
|
||
"Ideally the sequence $\\{\\mathbf{x}_k \\}_{k=0}$ converges to a global\n",
|
||
"minimum of the function $F$. In general we do not know if we are in a\n",
|
||
"global or local minimum. In the special case when $F$ is a convex\n",
|
||
"function, all local minima are also global minima, so in this case\n",
|
||
"gradient descent can converge to the global solution. The advantage of\n",
|
||
"this scheme is that it is conceptually simple and straightforward to\n",
|
||
"implement. However the method in this form has some severe\n",
|
||
"limitations:\n",
|
||
"\n",
|
||
"In machine learing we are often faced with non-convex high dimensional\n",
|
||
"cost functions with many local minima. Since GD is deterministic we\n",
|
||
"will get stuck in a local minimum, if the method converges, unless we\n",
|
||
"have a very good intial guess. This also implies that the scheme is\n",
|
||
"sensitive to the chosen initial condition.\n",
|
||
"\n",
|
||
"Note that the gradient is a function of $\\mathbf{x} =\n",
|
||
"(x_1,\\cdots,x_n)$ which makes it expensive to compute numerically.\n",
|
||
"\n",
|
||
"\n",
|
||
"<!-- !split -->\n",
|
||
"## The sensitiveness of the gradient descent\n",
|
||
"\n",
|
||
"The gradient descent method \n",
|
||
"is sensitive to the choice of learning rate $\\gamma_k$. This is due\n",
|
||
"to the fact that we are only guaranteed that $F(\\mathbf{x}_{k+1}) \\leq\n",
|
||
"F(\\mathbf{x}_k)$ for sufficiently small $\\gamma_k$. The problem is to\n",
|
||
"determine an optimal learning rate. If the learning rate is chosen too\n",
|
||
"small the method will take a long time to converge and if it is too\n",
|
||
"large we can experience erratic behavior.\n",
|
||
"\n",
|
||
"Many of these shortcomings can be alleviated by introducing\n",
|
||
"randomness. One such method is that of Stochastic Gradient Descent\n",
|
||
"(SGD), see below.\n",
|
||
"\n",
|
||
"\n",
|
||
"<!-- !split -->\n",
|
||
"## Convex functions\n",
|
||
"\n",
|
||
"Ideally we want our cost/loss function to be convex(concave).\n",
|
||
"\n",
|
||
"First we give the definition of a convex set: A set $C$ in\n",
|
||
"$\\mathbb{R}^n$ is said to be convex if, for all $x$ and $y$ in $C$ and\n",
|
||
"all $t \\in (0,1)$ , the point $(1 − t)x + ty$ also belongs to\n",
|
||
"C. Geometrically this means that every point on the line segment\n",
|
||
"connecting $x$ and $y$ is in $C$ as discussed below.\n",
|
||
"\n",
|
||
"The convex subsets of $\\mathbb{R}$ are the intervals of\n",
|
||
"$\\mathbb{R}$. Examples of convex sets of $\\mathbb{R}^2$ are the\n",
|
||
"regular polygons (triangles, rectangles, pentagons, etc...).\n",
|
||
"\n",
|
||
"## Convex function\n",
|
||
"\n",
|
||
"**Convex function**: Let $X \\subset \\mathbb{R}^n$ be a convex set. Assume that the function $f: X \\rightarrow \\mathbb{R}$ is continuous, then $f$ is said to be convex if $$f(tx_1 + (1-t)x_2) \\leq tf(x_1) + (1-t)f(x_2) $$ for all $x_1, x_2 \\in X$ and for all $t \\in [0,1]$. If $\\leq$ is replaced with a strict inequaltiy in the definition, we demand $x_1 \\neq x_2$ and $t\\in(0,1)$ then $f$ is said to be strictly convex. For a single variable function, convexity means that if you draw a straight line connecting $f(x_1)$ and $f(x_2)$, the value of the function on the interval $[x_1,x_2]$ is always below the line as illustrated below.\n",
|
||
"\n",
|
||
"## Conditions on convex functions\n",
|
||
"\n",
|
||
"In the following we state first and second-order conditions which\n",
|
||
"ensures convexity of a function $f$. We write $D_f$ to denote the\n",
|
||
"domain of $f$, i.e the subset of $R^n$ where $f$ is defined. For more\n",
|
||
"details and proofs we refer to: [S. Boyd and L. Vandenberghe. Convex Optimization. Cambridge University Press](http://stanford.edu/boyd/cvxbook/, 2004).\n",
|
||
"\n",
|
||
"**First order condition.**\n",
|
||
"\n",
|
||
"Suppose $f$ is differentiable (i.e $\\nabla f(x)$ is well defined for\n",
|
||
"all $x$ in the domain of $f$). Then $f$ is convex if and only if $D_f$\n",
|
||
"is a convex set and $$f(y) \\geq f(x) + \\nabla f(x)^T (y-x) $$ holds\n",
|
||
"for all $x,y \\in D_f$. This condition means that for a convex function\n",
|
||
"the first order Taylor expansion (right hand side above) at any point\n",
|
||
"a global under estimator of the function. To convince yourself you can\n",
|
||
"make a drawing of $f(x) = x^2+1$ and draw the tangent line to $f(x)$ and\n",
|
||
"note that it is always below the graph.\n",
|
||
"\n",
|
||
"\n",
|
||
"\n",
|
||
"**Second order condition.**\n",
|
||
"\n",
|
||
"Assume that $f$ is twice\n",
|
||
"differentiable, i.e the Hessian matrix exists at each point in\n",
|
||
"$D_f$. Then $f$ is convex if and only if $D_f$ is a convex set and its\n",
|
||
"Hessian is positive semi-definite for all $x\\in D_f$. For a\n",
|
||
"single-variable function this reduces to $f''(x) \\geq 0$. Geometrically this means that $f$ has nonnegative curvature\n",
|
||
"everywhere.\n",
|
||
"\n",
|
||
"\n",
|
||
"\n",
|
||
"This condition is particularly useful since it gives us an procedure for determining if the function under consideration is convex, apart from using the definition.\n",
|
||
"\n",
|
||
"## More on convex functions\n",
|
||
"\n",
|
||
"The next result is of great importance to us and the reason why we are\n",
|
||
"going on about convex functions. In machine learning we frequently\n",
|
||
"have to minimize a loss/cost function in order to find the best\n",
|
||
"parameters for the model we are considering. \n",
|
||
"\n",
|
||
"Ideally we want the\n",
|
||
"global minimum (for high-dimensional models it is hard to know\n",
|
||
"if we have local or global minimum). However, if the cost/loss function\n",
|
||
"is convex the following result provides invaluable information:\n",
|
||
"\n",
|
||
"**Any minimum is global for convex functions.**\n",
|
||
"\n",
|
||
"Consider the problem of finding $x \\in \\mathbb{R}^n$ such that $f(x)$\n",
|
||
"is minimal, where $f$ is convex and differentiable. Then, any point\n",
|
||
"$x^*$ that satisfies $\\nabla f(x^*) = 0$ is a global minimum.\n",
|
||
"\n",
|
||
"\n",
|
||
"\n",
|
||
"This result means that if we know that the cost/loss function is convex and we are able to find a minimum, we are guaranteed that it is a global minimum.\n",
|
||
"\n",
|
||
"## Some simple problems\n",
|
||
"\n",
|
||
"1. Show that $f(x)=x^2$ is convex for $x \\in \\mathbb{R}$ using the definition of convexity. Hint: If you re-write the definition, $f$ is convex if the following holds for all $x,y \\in D_f$ and any $\\lambda \\in [0,1]$ $\\lambda f(x)+(1-\\lambda)f(y)-f(\\lambda x + (1-\\lambda) y ) \\geq 0$.\n",
|
||
"\n",
|
||
"2. Using the second order condition show that the following functions are convex on the specified domain.\n",
|
||
"\n",
|
||
" * $f(x) = e^x$ is convex for $x \\in \\mathbb{R}$.\n",
|
||
"\n",
|
||
" * $g(x) = -\\ln(x)$ is convex for $x \\in (0,\\infty)$.\n",
|
||
"\n",
|
||
"\n",
|
||
"3. Let $f(x) = x^2$ and $g(x) = e^x$. Show that $f(g(x))$ and $g(f(x))$ is convex for $x \\in \\mathbb{R}$. Also show that if $f(x)$ is any convex function than $h(x) = e^{f(x)}$ is convex.\n",
|
||
"\n",
|
||
"4. A norm is any function that satisfy the following properties\n",
|
||
"\n",
|
||
" * $f(\\alpha x) = |\\alpha| f(x)$ for all $\\alpha \\in \\mathbb{R}$.\n",
|
||
"\n",
|
||
" * $f(x+y) \\leq f(x) + f(y)$\n",
|
||
"\n",
|
||
" * $f(x) \\leq 0$ for all $x \\in \\mathbb{R}^n$ with equality if and only if $x = 0$\n",
|
||
"\n",
|
||
"\n",
|
||
"Using the definition of convexity, try to show that a function satisfying the properties above is convex (the third condition is not needed to show this).\n",
|
||
"\n",
|
||
"\n",
|
||
"## Standard steepest descent\n",
|
||
"\n",
|
||
"\n",
|
||
"Before we proceed, we would like to discuss the approach called the\n",
|
||
"**standard Steepest descent**, which again leads to us having to be able\n",
|
||
"to compute a matrix. It belongs to the class of Conjugate Gradient methods (CG).\n",
|
||
"\n",
|
||
"[The success of the CG method](https://www.cs.cmu.edu/~quake-papers/painless-conjugate-gradient.pdf)\n",
|
||
"for finding solutions of non-linear problems is based on the theory\n",
|
||
"of conjugate gradients for linear systems of equations. It belongs to\n",
|
||
"the class of iterative methods for solving problems from linear\n",
|
||
"algebra of the type"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{A}\\hat{x} = \\hat{b}.\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"In the iterative process we end up with a problem like"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{r}= \\hat{b}-\\hat{A}\\hat{x},\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"where $\\hat{r}$ is the so-called residual or error in the iterative process.\n",
|
||
"\n",
|
||
"When we have found the exact solution, $\\hat{r}=0$.\n",
|
||
"\n",
|
||
"## Gradient method\n",
|
||
"\n",
|
||
"The residual is zero when we reach the minimum of the quadratic equation"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"P(\\hat{x})=\\frac{1}{2}\\hat{x}^T\\hat{A}\\hat{x} - \\hat{x}^T\\hat{b},\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"with the constraint that the matrix $\\hat{A}$ is positive definite and\n",
|
||
"symmetric. This defines also the Hessian and we want it to be positive definite. \n",
|
||
"\n",
|
||
"\n",
|
||
"## Steepest descent method\n",
|
||
"\n",
|
||
"We denote the initial guess for $\\hat{x}$ as $\\hat{x}_0$. \n",
|
||
"We can assume without loss of generality that"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{x}_0=0,\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"or consider the system"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{A}\\hat{z} = \\hat{b}-\\hat{A}\\hat{x}_0,\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"instead.\n",
|
||
"\n",
|
||
"\n",
|
||
"## Steepest descent method\n",
|
||
"One can show that the solution $\\hat{x}$ is also the unique minimizer of the quadratic form"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"f(\\hat{x}) = \\frac{1}{2}\\hat{x}^T\\hat{A}\\hat{x} - \\hat{x}^T \\hat{x} , \\quad \\hat{x}\\in\\mathbf{R}^n.\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"This suggests taking the first basis vector $\\hat{r}_1$ (see below for definition) \n",
|
||
"to be the gradient of $f$ at $\\hat{x}=\\hat{x}_0$, \n",
|
||
"which equals"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{A}\\hat{x}_0-\\hat{b},\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"and \n",
|
||
"$\\hat{x}_0=0$ it is equal $-\\hat{b}$.\n",
|
||
"\n",
|
||
"\n",
|
||
"\n",
|
||
"## Final expressions\n",
|
||
"We can compute the residual iteratively as"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{r}_{k+1}=\\hat{b}-\\hat{A}\\hat{x}_{k+1},\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"which equals"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{b}-\\hat{A}(\\hat{x}_k+\\alpha_k\\hat{r}_k),\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"or"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"(\\hat{b}-\\hat{A}\\hat{x}_k)-\\alpha_k\\hat{A}\\hat{r}_k,\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"which gives"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\alpha_k = \\frac{\\hat{r}_k^T\\hat{r}_k}{\\hat{r}_k^T\\hat{A}\\hat{r}_k}\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"leading to the iterative scheme"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{x}_{k+1}=\\hat{x}_k-\\alpha_k\\hat{r}_{k},\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"## Simple codes for steepest descent and conjugate gradient using a $2\\times 2$ matrix, in c++, Python code to come"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
" #include <cmath>\n",
|
||
" #include <iostream>\n",
|
||
" #include <fstream>\n",
|
||
" #include <iomanip>\n",
|
||
" #include \"vectormatrixclass.h\"\n",
|
||
" using namespace std;\n",
|
||
" // Main function begins here\n",
|
||
" int main(int argc, char * argv[]){\n",
|
||
" int dim = 2;\n",
|
||
" Vector x(dim),xsd(dim), b(dim),x0(dim);\n",
|
||
" Matrix A(dim,dim);\n",
|
||
" \n",
|
||
" // Set our initial guess\n",
|
||
" x0(0) = x0(1) = 0;\n",
|
||
" // Set the matrix\n",
|
||
" A(0,0) = 3; A(1,0) = 2; A(0,1) = 2; A(1,1) = 6;\n",
|
||
" b(0) = 2; b(1) = -8;\n",
|
||
" cout << \"The Matrix A that we are using: \" << endl;\n",
|
||
" A.Print();\n",
|
||
" cout << endl;\n",
|
||
" xsd = SteepestDescent(A,b,x0);\n",
|
||
" cout << \"The approximate solution using Steepest Descent is: \" << endl;\n",
|
||
" xsd.Print();\n",
|
||
" cout << endl;\n",
|
||
" }\n"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"## The routine for the steepest descent method"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
" Vector SteepestDescent(Matrix A, Vector b, Vector x0){\n",
|
||
" int IterMax, i;\n",
|
||
" int dim = x0.Dimension();\n",
|
||
" const double tolerance = 1.0e-14;\n",
|
||
" Vector x(dim),f(dim),z(dim);\n",
|
||
" double c,alpha,d;\n",
|
||
" IterMax = 30;\n",
|
||
" x = x0;\n",
|
||
" r = A*x-b;\n",
|
||
" i = 0;\n",
|
||
" while (i <= IterMax){\n",
|
||
" z = A*r;\n",
|
||
" c = dot(r,r);\n",
|
||
" alpha = c/dot(r,z);\n",
|
||
" x = x - alpha*r;\n",
|
||
" r = A*x-b;\n",
|
||
" if(sqrt(dot(r,r)) < tolerance) break;\n",
|
||
" i++;\n",
|
||
" }\n",
|
||
" return x;\n",
|
||
" }\n"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"## Steepest descent example"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "code",
|
||
"execution_count": 1,
|
||
"metadata": {
|
||
"collapsed": false
|
||
},
|
||
"outputs": [],
|
||
"source": [
|
||
"%matplotlib inline\n",
|
||
"\n",
|
||
"import numpy as np\n",
|
||
"import numpy.linalg as la\n",
|
||
"\n",
|
||
"import scipy.optimize as sopt\n",
|
||
"\n",
|
||
"import matplotlib.pyplot as pt\n",
|
||
"from mpl_toolkits.mplot3d import axes3d\n",
|
||
"\n",
|
||
"def f(x):\n",
|
||
" return 0.5*x[0]**2 + 2.5*x[1]**2\n",
|
||
"\n",
|
||
"def df(x):\n",
|
||
" return np.array([x[0], 5*x[1]])\n",
|
||
"\n",
|
||
"fig = pt.figure()\n",
|
||
"ax = fig.gca(projection=\"3d\")\n",
|
||
"\n",
|
||
"xmesh, ymesh = np.mgrid[-2:2:50j,-2:2:50j]\n",
|
||
"fmesh = f(np.array([xmesh, ymesh]))\n",
|
||
"ax.plot_surface(xmesh, ymesh, fmesh)"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"And then as countor plot"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "code",
|
||
"execution_count": 2,
|
||
"metadata": {
|
||
"collapsed": false
|
||
},
|
||
"outputs": [],
|
||
"source": [
|
||
"pt.axis(\"equal\")\n",
|
||
"pt.contour(xmesh, ymesh, fmesh)\n",
|
||
"guesses = [np.array([2, 2./5])]"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"Find guesses"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "code",
|
||
"execution_count": 3,
|
||
"metadata": {
|
||
"collapsed": false
|
||
},
|
||
"outputs": [],
|
||
"source": [
|
||
"x = guesses[-1]\n",
|
||
"s = -df(x)"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"Run it!"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "code",
|
||
"execution_count": 4,
|
||
"metadata": {
|
||
"collapsed": false
|
||
},
|
||
"outputs": [],
|
||
"source": [
|
||
"def f1d(alpha):\n",
|
||
" return f(x + alpha*s)\n",
|
||
"\n",
|
||
"alpha_opt = sopt.golden(f1d)\n",
|
||
"next_guess = x + alpha_opt * s\n",
|
||
"guesses.append(next_guess)\n",
|
||
"print(next_guess)"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"What happened?"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "code",
|
||
"execution_count": 5,
|
||
"metadata": {
|
||
"collapsed": false
|
||
},
|
||
"outputs": [],
|
||
"source": [
|
||
"pt.axis(\"equal\")\n",
|
||
"pt.contour(xmesh, ymesh, fmesh, 50)\n",
|
||
"it_array = np.array(guesses)\n",
|
||
"pt.plot(it_array.T[0], it_array.T[1], \"x-\")"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"## Conjugate gradient\n",
|
||
"\n",
|
||
"\n",
|
||
"\n",
|
||
"<!-- !split -->\n",
|
||
"## Revisiting our first homework\n",
|
||
"\n",
|
||
"We will use linear regression as a case study for the gradient descent\n",
|
||
"methods. Linear regression is a great test case for the gradient\n",
|
||
"descent methods discussed in the lectures since it has several\n",
|
||
"desirable properties such as:\n",
|
||
"\n",
|
||
"1. An analytical solution (recall homework set 1).\n",
|
||
"\n",
|
||
"2. The gradient can be computed analytically.\n",
|
||
"\n",
|
||
"3. The cost function is convex which guarantees that gradient descent converges for small enough learning rates\n",
|
||
"\n",
|
||
"We revisit the example from homework set 1 where we had"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"y_i = 5x_i^2 + 0.1\\xi_i, \\ i=1,\\cdots,100\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"with $x_i \\in [0,1] $ chosen randomly with a uniform distribution. Additionally $\\xi_i$ represents stochastic noise chosen according to a normal distribution $\\cal {N}(0,1)$. \n",
|
||
"The linear regression model is given by"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"h_\\beta(x) = \\hat{y} = \\beta_0 + \\beta_1 x,\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"such that"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{y}_i = \\beta_0 + \\beta_1 x_i.\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"<!-- !split -->\n",
|
||
"## Gradient descent example\n",
|
||
"\n",
|
||
"Let $\\mathbf{y} = (y_1,\\cdots,y_n)^T$, $\\mathbf{\\hat{y}} = (\\hat{y}_1,\\cdots,\\hat{y}_n)^T$ and $\\beta = (\\beta_0, \\beta_1)^T$\n",
|
||
"\n",
|
||
"It is convenient to write $\\mathbf{\\hat{y}} = X\\beta$ where $X \\in \\mathbb{R}^{100 \\times 2} $ is the design matrix given by"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"X \\equiv \\begin{bmatrix}\n",
|
||
"1 & x_1 \\\\\n",
|
||
"\\vdots & \\vdots \\\\\n",
|
||
"1 & x_{100} & \\\\\n",
|
||
"\\end{bmatrix}.\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"The loss function is given by"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"C(\\beta) = ||X\\beta-\\mathbf{y}||^2 = ||X\\beta||^2 - 2 \\mathbf{y}^T X\\beta + ||\\mathbf{y}||^2 = \\sum_{i=1}^{100} (\\beta_0 + \\beta_1 x_i)^2 - 2 y_i (\\beta_0 + \\beta_1 x_i) + y_i^2\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"and we want to find $\\beta$ such that $C(\\beta)$ is minimized.\n",
|
||
"\n",
|
||
"## The derivative of the cost/loss function\n",
|
||
"\n",
|
||
"Computing $\\partial C(\\beta) / \\partial \\beta_0$ and $\\partial C(\\beta) / \\partial \\beta_1$ we can show that the gradient can be written as"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\nabla_{\\beta} C(\\beta) = (\\partial C(\\beta) / \\partial \\beta_0, \\partial C(\\beta) / \\partial \\beta_1)^T = 2\\begin{bmatrix} \\sum_{i=1}^{100} \\left(\\beta_0+\\beta_1x_i-y_i\\right) \\\\\n",
|
||
"\\sum_{i=1}^{100}\\left( x_i (\\beta_0+\\beta_1x_i)-y_ix_i\\right) \\\\\n",
|
||
"\\end{bmatrix} = 2X^T(X\\beta - \\mathbf{y}),\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"where $X$ is the design matrix defined above.\n",
|
||
"\n",
|
||
"## The Hessian matrix\n",
|
||
"The Hessian matrix of $C(\\beta)$ is given by"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{H} \\equiv \\begin{bmatrix}\n",
|
||
"\\frac{\\partial^2 C(\\beta)}{\\partial \\beta_0^2} & \\frac{\\partial^2 C(\\beta)}{\\partial \\beta_0 \\partial \\beta_1} \\\\\n",
|
||
"\\frac{\\partial^2 C(\\beta)}{\\partial \\beta_0 \\partial \\beta_1} & \\frac{\\partial^2 C(\\beta)}{\\partial \\beta_1^2} & \\\\\n",
|
||
"\\end{bmatrix} = 2X^T X.\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"This result implies that $C(\\beta)$ is a convex function since the matrix $X^T X$ always is positive semi-definite.\n",
|
||
"\n",
|
||
"\n",
|
||
"\n",
|
||
"\n",
|
||
"## Simple program\n",
|
||
"\n",
|
||
"We can now write a program that minimizes $C(\\beta)$ using the gradient descent method with a constant learning rate $\\gamma$ according to"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\beta_{k+1} = \\beta_k - \\gamma \\nabla_\\beta C(\\beta_k), \\ k=0,1,\\cdots\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"We can use the expression we computed for the gradient and let use a\n",
|
||
"$\\beta_0$ be chosen randomly and let $\\gamma = 0.001$. Stop iterating\n",
|
||
"when $||\\nabla_\\beta C(\\beta_k) || \\leq \\epsilon = 10^{-8}$. \n",
|
||
"\n",
|
||
"And finally we can compare our solution for $\\beta$ with the analytic result given by \n",
|
||
"$\\beta= (X^TX)^{-1} X^T \\mathbf{y}$."
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "code",
|
||
"execution_count": 6,
|
||
"metadata": {
|
||
"collapsed": false
|
||
},
|
||
"outputs": [],
|
||
"source": [
|
||
"import numpy as np\n",
|
||
"\n",
|
||
"\"\"\"\n",
|
||
"The following setup is just a suggestion, feel free to write it the way you like.\n",
|
||
"\"\"\"\n",
|
||
"\n",
|
||
"#Setup problem described in the exercise\n",
|
||
"N = 100 #Nr of datapoints\n",
|
||
"M = 2 #Nr of features\n",
|
||
"x = np.random.rand(N) #Uniformly generated x-values in [0,1]\n",
|
||
"y = 5*x**2 + 0.1*np.random.randn(N)\n",
|
||
"X = np.c_[np.ones(N),x] #Construct design matrix\n",
|
||
"\n",
|
||
"#Compute beta according to normal equations to compare with GD solution\n",
|
||
"Xt_X_inv = np.linalg.inv(np.dot(X.T,X))\n",
|
||
"Xt_y = np.dot(X.transpose(),y)\n",
|
||
"beta_NE = np.dot(Xt_X_inv,Xt_y)\n",
|
||
"print(beta_NE)"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"## Gradient Descent Example\n",
|
||
"\n",
|
||
"Another simple example is here"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "code",
|
||
"execution_count": 7,
|
||
"metadata": {
|
||
"collapsed": false
|
||
},
|
||
"outputs": [],
|
||
"source": [
|
||
"\n",
|
||
"# Importing various packages\n",
|
||
"from random import random, seed\n",
|
||
"import numpy as np\n",
|
||
"import matplotlib.pyplot as plt\n",
|
||
"from mpl_toolkits.mplot3d import Axes3D\n",
|
||
"from matplotlib import cm\n",
|
||
"from matplotlib.ticker import LinearLocator, FormatStrFormatter\n",
|
||
"import sys\n",
|
||
"\n",
|
||
"x = 2*np.random.rand(100,1)\n",
|
||
"y = 4+3*x+np.random.randn(100,1)\n",
|
||
"\n",
|
||
"xb = np.c_[np.ones((100,1)), x]\n",
|
||
"beta_linreg = np.linalg.inv(xb.T.dot(xb)).dot(xb.T).dot(y)\n",
|
||
"print(beta_linreg)\n",
|
||
"beta = np.random.randn(2,1)\n",
|
||
"\n",
|
||
"eta = 0.1\n",
|
||
"Niterations = 1000\n",
|
||
"m = 100\n",
|
||
"\n",
|
||
"for iter in range(Niterations):\n",
|
||
" gradients = 2.0/m*xb.T.dot(xb.dot(beta)-y)\n",
|
||
" beta -= eta*gradients\n",
|
||
"\n",
|
||
"print(beta)\n",
|
||
"xnew = np.array([[0],[2]])\n",
|
||
"xbnew = np.c_[np.ones((2,1)), xnew]\n",
|
||
"ypredict = xbnew.dot(beta)\n",
|
||
"ypredict2 = xbnew.dot(beta_linreg)\n",
|
||
"plt.plot(xnew, ypredict, \"r-\")\n",
|
||
"plt.plot(xnew, ypredict2, \"b-\")\n",
|
||
"plt.plot(x, y ,'ro')\n",
|
||
"plt.axis([0,2.0,0, 15.0])\n",
|
||
"plt.xlabel(r'$x$')\n",
|
||
"plt.ylabel(r'$y$')\n",
|
||
"plt.title(r'Gradient descent example')\n",
|
||
"plt.show()"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"## And a corresponding example using **scikit-learn**"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "code",
|
||
"execution_count": 8,
|
||
"metadata": {
|
||
"collapsed": false
|
||
},
|
||
"outputs": [],
|
||
"source": [
|
||
"# Importing various packages\n",
|
||
"from random import random, seed\n",
|
||
"import numpy as np\n",
|
||
"import matplotlib.pyplot as plt\n",
|
||
"from sklearn.linear_model import SGDRegressor\n",
|
||
"\n",
|
||
"x = 2*np.random.rand(100,1)\n",
|
||
"y = 4+3*x+np.random.randn(100,1)\n",
|
||
"\n",
|
||
"xb = np.c_[np.ones((100,1)), x]\n",
|
||
"beta_linreg = np.linalg.inv(xb.T.dot(xb)).dot(xb.T).dot(y)\n",
|
||
"print(beta_linreg)\n",
|
||
"sgdreg = SGDRegressor(n_iter = 50, penalty=None, eta0=0.1)\n",
|
||
"sgdreg.fit(x,y.ravel())\n",
|
||
"print(sgdreg.intercept_, sgdreg.coef_)"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"<!-- !split -->\n",
|
||
"## Gradient descent and Ridge\n",
|
||
"\n",
|
||
"We have also discussed Ridge regression where the loss function contains a regularized given by the $L_2$ norm of $\\beta$,"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"C_{\\text{ridge}}(\\beta) = ||X\\beta -\\mathbf{y}||^2 + \\lambda ||\\beta||^2, \\ \\lambda \\geq 0.\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"In order to minimize $C_{\\text{ridge}}(\\beta)$ using GD we only have adjust the gradient as follows"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\nabla_\\beta C_{\\text{ridge}}(\\beta) = 2\\begin{bmatrix} \\sum_{i=1}^{100} \\left(\\beta_0+\\beta_1x_i-y_i\\right) \\\\\n",
|
||
"\\sum_{i=1}^{100}\\left( x_i (\\beta_0+\\beta_1x_i)-y_ix_i\\right) \\\\\n",
|
||
"\\end{bmatrix} + 2\\lambda\\begin{bmatrix} \\beta_0 \\\\ \\beta_1\\end{bmatrix} = 2 (X^T(X\\beta - \\mathbf{y})+\\lambda \\beta).\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"We can now extend our program to minimize $C_{\\text{ridge}}(\\beta)$ using gradient descent and compare with the analytical solution given by"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\beta_{\\text{ridge}} = \\left(X^T X + \\lambda I_{2 \\times 2} \\right)^{-1} X^T \\mathbf{y},\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"for $\\lambda = {0,1,10,50,100}$ ($\\lambda = 0$ corresponds to ordinary least squares). \n",
|
||
"We can then compute $||\\beta_{\\text{ridge}}||$ for each $\\lambda$."
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "code",
|
||
"execution_count": 9,
|
||
"metadata": {
|
||
"collapsed": false
|
||
},
|
||
"outputs": [],
|
||
"source": [
|
||
"import numpy as np\n",
|
||
"\n",
|
||
"\"\"\"\n",
|
||
"The following setup is just a suggestion, feel free to write it the way you like.\n",
|
||
"\"\"\"\n",
|
||
"\n",
|
||
"#Setup problem described in the exercise\n",
|
||
"N = 100 #Nr of datapoints\n",
|
||
"M = 2 #Nr of features\n",
|
||
"x = np.random.rand(N)\n",
|
||
"y = 5*x**2 + 0.1*np.random.randn(N)\n",
|
||
"\n",
|
||
"\n",
|
||
"#Compute analytic beta for Ridge regression \n",
|
||
"X = np.c_[np.ones(N),x]\n",
|
||
"XT_X = np.dot(X.T,X)\n",
|
||
"\n",
|
||
"l = 0.1 #Ridge parameter lambda\n",
|
||
"Id = np.eye(XT_X.shape[0])\n",
|
||
"\n",
|
||
"Z = np.linalg.inv(XT_X+l*Id)\n",
|
||
"beta_ridge = np.dot(Z,np.dot(X.T,y))\n",
|
||
"\n",
|
||
"print(beta_ridge)\n",
|
||
"print(np.linalg.norm(beta_ridge)) #||beta||"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"## Automatic differentiation\n",
|
||
"Python has tools for so-called **automatic differentiation**.\n",
|
||
"Consider the following example"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"f(x) = \\sin\\left(2\\pi x + x^2\\right)\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"which has the following derivative"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"f'(x) = \\cos\\left(2\\pi x + x^2\\right)\\left(2\\pi + 2x\\right)\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"Using **autograd** we have"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "code",
|
||
"execution_count": 10,
|
||
"metadata": {
|
||
"collapsed": false
|
||
},
|
||
"outputs": [],
|
||
"source": [
|
||
"import autograd.numpy as np\n",
|
||
"\n",
|
||
"# To do elementwise differentiation:\n",
|
||
"from autograd import elementwise_grad as egrad \n",
|
||
"\n",
|
||
"# To plot:\n",
|
||
"import matplotlib.pyplot as plt \n",
|
||
"\n",
|
||
"\n",
|
||
"def f(x):\n",
|
||
" return np.sin(2*np.pi*x + x**2)\n",
|
||
"\n",
|
||
"def f_grad_analytic(x):\n",
|
||
" return np.cos(2*np.pi*x + x**2)*(2*np.pi + 2*x)\n",
|
||
"\n",
|
||
"# Do the comparison:\n",
|
||
"x = np.linspace(0,1,1000)\n",
|
||
"\n",
|
||
"f_grad = egrad(f)\n",
|
||
"\n",
|
||
"computed = f_grad(x)\n",
|
||
"analytic = f_grad_analytic(x)\n",
|
||
"\n",
|
||
"plt.title('Derivative computed from Autograd compared with the analytical derivative')\n",
|
||
"plt.plot(x,computed,label='autograd')\n",
|
||
"plt.plot(x,analytic,label='analytic')\n",
|
||
"\n",
|
||
"plt.xlabel('x')\n",
|
||
"plt.ylabel('y')\n",
|
||
"plt.legend()\n",
|
||
"\n",
|
||
"plt.show()\n",
|
||
"\n",
|
||
"print(\"The max absolute difference is: %g\"%(np.max(np.abs(computed - analytic))))"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"<!-- !split -->\n",
|
||
"## Using autograd\n",
|
||
"\n",
|
||
"Here we\n",
|
||
"experiment with what kind of functions Autograd is capable\n",
|
||
"of finding the gradient of. The following Python functions are just\n",
|
||
"meant to illustrate what Autograd can do, but please feel free to\n",
|
||
"experiment with other, possibly more complicated, functions as well."
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "code",
|
||
"execution_count": 11,
|
||
"metadata": {
|
||
"collapsed": false
|
||
},
|
||
"outputs": [],
|
||
"source": [
|
||
"import autograd.numpy as np\n",
|
||
"from autograd import grad\n",
|
||
"\n",
|
||
"def f1(x):\n",
|
||
" return x**3 + 1\n",
|
||
"\n",
|
||
"f1_grad = grad(f1)\n",
|
||
"\n",
|
||
"# Remember to send in float as argument to the computed gradient from Autograd!\n",
|
||
"a = 1.0\n",
|
||
"\n",
|
||
"# See the evaluated gradient at a using autograd:\n",
|
||
"print(\"The gradient of f1 evaluated at a = %g using autograd is: %g\"%(a,f1_grad(a)))\n",
|
||
"\n",
|
||
"# Compare with the analytical derivative, that is f1'(x) = 3*x**2 \n",
|
||
"grad_analytical = 3*a**2\n",
|
||
"print(\"The gradient of f1 evaluated at a = %g by finding the analytic expression is: %g\"%(a,grad_analytical))"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"## Autograd with more complicated functions\n",
|
||
"\n",
|
||
"To differentiate with respect to two (or more) arguments of a Python\n",
|
||
"function, Autograd need to know at which variable the function if\n",
|
||
"being differentiated with respect to."
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "code",
|
||
"execution_count": 12,
|
||
"metadata": {
|
||
"collapsed": false
|
||
},
|
||
"outputs": [],
|
||
"source": [
|
||
"import autograd.numpy as np\n",
|
||
"from autograd import grad\n",
|
||
"def f2(x1,x2):\n",
|
||
" return 3*x1**3 + x2*(x1 - 5) + 1\n",
|
||
"\n",
|
||
"# By sending the argument 0, Autograd will compute the derivative w.r.t the first variable, in this case x1\n",
|
||
"f2_grad_x1 = grad(f2,0)\n",
|
||
"\n",
|
||
"# ... and differentiate w.r.t x2 by sending 1 as an additional arugment to grad\n",
|
||
"f2_grad_x2 = grad(f2,1)\n",
|
||
"\n",
|
||
"x1 = 1.0\n",
|
||
"x2 = 3.0 \n",
|
||
"\n",
|
||
"print(\"Evaluating at x1 = %g, x2 = %g\"%(x1,x2))\n",
|
||
"print(\"-\"*30)\n",
|
||
"\n",
|
||
"# Compare with the analytical derivatives:\n",
|
||
"\n",
|
||
"# Derivative of f2 w.r.t x1 is: 9*x1**2 + x2:\n",
|
||
"f2_grad_x1_analytical = 9*x1**2 + x2\n",
|
||
"\n",
|
||
"# Derivative of f2 w.r.t x2 is: x1 - 5:\n",
|
||
"f2_grad_x2_analytical = x1 - 5\n",
|
||
"\n",
|
||
"# See the evaluated derivations:\n",
|
||
"print(\"The derivative of f2 w.r.t x1: %g\"%( f2_grad_x1(x1,x2) ))\n",
|
||
"print(\"The analytical derivative of f2 w.r.t x1: %g\"%( f2_grad_x1(x1,x2) ))\n",
|
||
"\n",
|
||
"print()\n",
|
||
"\n",
|
||
"print(\"The derivative of f2 w.r.t x2: %g\"%( f2_grad_x2(x1,x2) ))\n",
|
||
"print(\"The analytical derivative of f2 w.r.t x2: %g\"%( f2_grad_x2(x1,x2) ))"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"Note that the grad function will not produce the true gradient of the function. The true gradient of a function with two or more variables will produce a vector, where each element is the function differentiated w.r.t a variable.\n",
|
||
"\n",
|
||
"\n",
|
||
"## More complicated functions using the elements of their arguments directly"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "code",
|
||
"execution_count": 13,
|
||
"metadata": {
|
||
"collapsed": false
|
||
},
|
||
"outputs": [],
|
||
"source": [
|
||
"import autograd.numpy as np\n",
|
||
"from autograd import grad\n",
|
||
"def f3(x): # Assumes x is an array of length 5 or higher\n",
|
||
" return 2*x[0] + 3*x[1] + 5*x[2] + 7*x[3] + 11*x[4]**2\n",
|
||
"\n",
|
||
"f3_grad = grad(f3)\n",
|
||
"\n",
|
||
"x = np.linspace(0,4,5)\n",
|
||
"\n",
|
||
"# Print the computed gradient:\n",
|
||
"print(\"The computed gradient of f3 is: \", f3_grad(x))\n",
|
||
"\n",
|
||
"# The analytical gradient is: (2, 3, 5, 7, 22*x[4])\n",
|
||
"f3_grad_analytical = np.array([2, 3, 5, 7, 22*x[4]])\n",
|
||
"\n",
|
||
"# Print the analytical gradient:\n",
|
||
"print(\"The analytical gradient of f3 is: \", f3_grad_analytical)"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"Note that in this case, when sending an array as input argument, the\n",
|
||
"output from Autograd is another array. This is the true gradient of\n",
|
||
"the function, as opposed to the function in the previous example. By\n",
|
||
"using arrays to represent the variables, the output from Autograd\n",
|
||
"might be easier to work with, as the output is closer to what one\n",
|
||
"could expect form a gradient-evaluting function.\n",
|
||
"\n",
|
||
"<!-- !split -->\n",
|
||
"## Functions using mathematical functions from Numpy"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "code",
|
||
"execution_count": 14,
|
||
"metadata": {
|
||
"collapsed": false
|
||
},
|
||
"outputs": [],
|
||
"source": [
|
||
"import autograd.numpy as np\n",
|
||
"from autograd import grad\n",
|
||
"def f4(x):\n",
|
||
" return np.sqrt(1+x**2) + np.exp(x) + np.sin(2*np.pi*x)\n",
|
||
"\n",
|
||
"f4_grad = grad(f4)\n",
|
||
"\n",
|
||
"x = 2.7\n",
|
||
"\n",
|
||
"# Print the computed derivative:\n",
|
||
"print(\"The computed derivative of f4 at x = %g is: %g\"%(x,f4_grad(x)))\n",
|
||
"\n",
|
||
"# The analytical derivative is: x/sqrt(1 + x**2) + exp(x) + cos(2*pi*x)*2*pi\n",
|
||
"f4_grad_analytical = x/np.sqrt(1 + x**2) + np.exp(x) + np.cos(2*np.pi*x)*2*np.pi\n",
|
||
"\n",
|
||
"# Print the analytical gradient:\n",
|
||
"print(\"The analytical gradient of f4 at x = %g is: %g\"%(x,f4_grad_analytical))"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"## More autograd"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "code",
|
||
"execution_count": 15,
|
||
"metadata": {
|
||
"collapsed": false
|
||
},
|
||
"outputs": [],
|
||
"source": [
|
||
"import autograd.numpy as np\n",
|
||
"from autograd import grad\n",
|
||
"def f5(x):\n",
|
||
" if x >= 0:\n",
|
||
" return x**2\n",
|
||
" else:\n",
|
||
" return -3*x + 1\n",
|
||
"\n",
|
||
"f5_grad = grad(f5)\n",
|
||
"\n",
|
||
"x = 2.7\n",
|
||
"\n",
|
||
"# Print the computed derivative:\n",
|
||
"print(\"The computed derivative of f5 at x = %g is: %g\"%(x,f5_grad(x)))"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"## And with loops"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"1\n",
|
||
"7\n",
|
||
" \n",
|
||
"<\n",
|
||
"<\n",
|
||
"<\n",
|
||
"!\n",
|
||
"!\n",
|
||
"C\n",
|
||
"O\n",
|
||
"D\n",
|
||
"E\n",
|
||
"_\n",
|
||
"B\n",
|
||
"L\n",
|
||
"O\n",
|
||
"C\n",
|
||
"K\n",
|
||
" \n",
|
||
" \n",
|
||
"p\n",
|
||
"y\n",
|
||
"c\n",
|
||
"o\n",
|
||
"d"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "code",
|
||
"execution_count": 16,
|
||
"metadata": {
|
||
"collapsed": false
|
||
},
|
||
"outputs": [],
|
||
"source": [
|
||
"import autograd.numpy as np\n",
|
||
"from autograd import grad\n",
|
||
"# Both of the functions are implementation of the sum: sum(x**i) for i = 0, ..., 9\n",
|
||
"# The analytical derivative is: sum(i*x**(i-1)) \n",
|
||
"f6_grad_analytical = 0\n",
|
||
"for i in range(10):\n",
|
||
" f6_grad_analytical += i*x**(i-1)\n",
|
||
"\n",
|
||
"print(\"The analytical derivative of f6 at x = %g is: %g\"%(x,f6_grad_analytical))"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"## Using recursion"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "code",
|
||
"execution_count": 17,
|
||
"metadata": {
|
||
"collapsed": false
|
||
},
|
||
"outputs": [],
|
||
"source": [
|
||
"import autograd.numpy as np\n",
|
||
"from autograd import grad\n",
|
||
"\n",
|
||
"def f7(n): # Assume that n is an integer\n",
|
||
" if n == 1 or n == 0:\n",
|
||
" return 1\n",
|
||
" else:\n",
|
||
" return n*f7(n-1)\n",
|
||
"\n",
|
||
"f7_grad = grad(f7)\n",
|
||
"\n",
|
||
"n = 2.0\n",
|
||
"\n",
|
||
"print(\"The computed derivative of f7 at n = %d is: %g\"%(n,f7_grad(n)))\n",
|
||
"\n",
|
||
"# The function f7 is an implementation of the factorial of n.\n",
|
||
"# By using the product rule, one can find that the derivative is:\n",
|
||
"\n",
|
||
"f7_grad_analytical = 0\n",
|
||
"for i in range(int(n)-1):\n",
|
||
" tmp = 1\n",
|
||
" for k in range(int(n)-1):\n",
|
||
" if k != i:\n",
|
||
" tmp *= (n - k)\n",
|
||
" f7_grad_analytical += tmp\n",
|
||
"\n",
|
||
"print(\"The analytical derivative of f7 at n = %d is: %g\"%(n,f7_grad_analytical))"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"Note that if n is equal to zero or one, Autograd will give an error message. This message appears when the output is independent on input.\n",
|
||
"\n",
|
||
"## Unsupported functions\n",
|
||
"Autograd supports many features. However, there are some functions that is not supported (yet) by Autograd.\n",
|
||
"\n",
|
||
"Assigning a value to the variable being differentiated with respect to"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "code",
|
||
"execution_count": 18,
|
||
"metadata": {
|
||
"collapsed": false
|
||
},
|
||
"outputs": [],
|
||
"source": [
|
||
"import autograd.numpy as np\n",
|
||
"from autograd import grad\n",
|
||
"def f8(x): # Assume x is an array\n",
|
||
" x[2] = 3\n",
|
||
" return x*2\n",
|
||
"\n",
|
||
"f8_grad = grad(f8)\n",
|
||
"\n",
|
||
"x = 8.4\n",
|
||
"\n",
|
||
"print(\"The derivative of f8 is:\",f8_grad(x))"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"Here, Autograd tells us that an 'ArrayBox' does not support item assignment. The item assignment is done when the program tries to assign x[2] to the value 3. However, Autograd has implemented the computation of the derivative such that this assignment is not possible.\n",
|
||
"\n",
|
||
"## The syntax a.dot(b) when finding the dot product"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "code",
|
||
"execution_count": 19,
|
||
"metadata": {
|
||
"collapsed": false
|
||
},
|
||
"outputs": [],
|
||
"source": [
|
||
"import autograd.numpy as np\n",
|
||
"from autograd import grad\n",
|
||
"def f9(a): # Assume a is an array with 2 elements\n",
|
||
" b = np.array([1.0,2.0])\n",
|
||
" return a.dot(b)\n",
|
||
"\n",
|
||
"f9_grad = grad(f9)\n",
|
||
"\n",
|
||
"x = np.array([1.0,0.0])\n",
|
||
"\n",
|
||
"print(\"The derivative of f9 is:\",f9_grad(x))"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"Here we are told that the 'dot' function does not belong to Autograd's\n",
|
||
"version of a Numpy array. To overcome this, an alternative syntax\n",
|
||
"which also computed the dot product can be used:"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "code",
|
||
"execution_count": 20,
|
||
"metadata": {
|
||
"collapsed": false
|
||
},
|
||
"outputs": [],
|
||
"source": [
|
||
"import autograd.numpy as np\n",
|
||
"from autograd import grad\n",
|
||
"def f9_alternative(x): # Assume a is an array with 2 elements\n",
|
||
" b = np.array([1.0,2.0])\n",
|
||
" return np.dot(x,b) # The same as x_1*b_1 + x_2*b_2\n",
|
||
"\n",
|
||
"f9_alternative_grad = grad(f9_alternative)\n",
|
||
"\n",
|
||
"x = np.array([3.0,0.0])\n",
|
||
"\n",
|
||
"print(\"The gradient of f9 is:\",f9_alternative_grad(x))\n",
|
||
"\n",
|
||
"# The analytical gradient of the dot product of vectors x and b with two elements (x_1,x_2) and (b_1, b_2) respectively\n",
|
||
"# w.r.t x is (b_1, b_2)."
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"## Recommended to avoid\n",
|
||
"The documentation recommends to avoid inplace operations such as"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "code",
|
||
"execution_count": 21,
|
||
"metadata": {
|
||
"collapsed": false
|
||
},
|
||
"outputs": [],
|
||
"source": [
|
||
"a += b\n",
|
||
"a -= b\n",
|
||
"a*= b\n",
|
||
"a /=b"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"## Stochastic Gradient Descent\n",
|
||
"\n",
|
||
"Stochastic gradient descent (SGD) and variants thereof address some of\n",
|
||
"the shortcomings of the Gradient descent method discussed above.\n",
|
||
"\n",
|
||
"The underlying idea of SGD comes from the observation that the cost\n",
|
||
"function, which we want to minimize, can almost always be written as a\n",
|
||
"sum over $n$ data points $\\{\\mathbf{x}_i\\}_{i=1}^n$,"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"C(\\mathbf{\\beta}) = \\sum_{i=1}^n c_i(\\mathbf{x}_i,\n",
|
||
"\\mathbf{\\beta}).\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"## Computation of gradients\n",
|
||
"\n",
|
||
"This in turn means that the gradient can be\n",
|
||
"computed as a sum over $i$-gradients"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\nabla_\\beta C(\\mathbf{\\beta}) = \\sum_i^n \\nabla_\\beta c_i(\\mathbf{x}_i,\n",
|
||
"\\mathbf{\\beta}).\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"Stochasticity/randomness is introduced by only taking the\n",
|
||
"gradient on a subset of the data called minibatches. If there are $n$\n",
|
||
"data points and the size of each minibatch is $M$, there will be $n/M$\n",
|
||
"minibatches. We denote these minibatches by $B_k$ where\n",
|
||
"$k=1,\\cdots,n/M$.\n",
|
||
"\n",
|
||
"## SGD example\n",
|
||
"As an example, suppose we have $10$ data points $(\\mathbf{x}_1,\\cdots, \\mathbf{x}_{10})$ \n",
|
||
"and we choose to have $M=5$ minibathces,\n",
|
||
"then each minibatch contains two data points. In particular we have\n",
|
||
"$B_1 = (\\mathbf{x}_1,\\mathbf{x}_2), \\cdots, B_5 =\n",
|
||
"(\\mathbf{x}_9,\\mathbf{x}_{10})$. Note that if you choose $M=1$ you\n",
|
||
"have only a single batch with all data points and on the other extreme,\n",
|
||
"you may choose $M=n$ resulting in a minibatch for each datapoint, i.e\n",
|
||
"$B_k = \\mathbf{x}_k$.\n",
|
||
"\n",
|
||
"The idea is now to approximate the gradient by replacing the sum over\n",
|
||
"all data points with a sum over the data points in one the minibatches\n",
|
||
"picked at random in each gradient descent step"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\nabla_{\\beta}\n",
|
||
"C(\\mathbf{\\beta}) = \\sum_{i=1}^n \\nabla_\\beta c_i(\\mathbf{x}_i,\n",
|
||
"\\mathbf{\\beta}) \\rightarrow \\sum_{i \\in B_k}^n \\nabla_\\beta\n",
|
||
"c_i(\\mathbf{x}_i, \\mathbf{\\beta}).\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"## The gradient step\n",
|
||
"\n",
|
||
"Thus a gradient descent step now looks like"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\beta_{j+1} = \\beta_j - \\gamma_j \\sum_{i \\in B_k}^n \\nabla_\\beta c_i(\\mathbf{x}_i,\n",
|
||
"\\mathbf{\\beta})\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"where $k$ is picked at random with equal\n",
|
||
"probability from $[1,n/M]$. An iteration over the number of\n",
|
||
"minibathces (n/M) is commonly referred to as an epoch. Thus it is\n",
|
||
"typical to choose a number of epochs and for each epoch iterate over\n",
|
||
"the number of minibatches, as exemplified in the code below.\n",
|
||
"\n",
|
||
"## Simple example code"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "code",
|
||
"execution_count": 22,
|
||
"metadata": {
|
||
"collapsed": false
|
||
},
|
||
"outputs": [],
|
||
"source": [
|
||
"import numpy as np \n",
|
||
"\n",
|
||
"n = 100 #100 datapoints \n",
|
||
"M = 5 #size of each minibatch\n",
|
||
"m = int(n/M) #number of minibatches\n",
|
||
"n_epochs = 10 #number of epochs\n",
|
||
"\n",
|
||
"j = 0\n",
|
||
"for epoch in range(1,n_epochs+1):\n",
|
||
" for i in range(m):\n",
|
||
" k = np.random.randint(m) #Pick the k-th minibatch at random\n",
|
||
" #Compute the gradient using the data in minibatch Bk\n",
|
||
" #Compute new suggestion for \n",
|
||
" j += 1"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"Taking the gradient only on a subset of the data has two important\n",
|
||
"benefits. First, it introduces randomness which decreases the chance\n",
|
||
"that our opmization scheme gets stuck in a local minima. Second, if\n",
|
||
"the size of the minibatches are small relative to the number of\n",
|
||
"datapoints ($M < n$), the computation of the gradient is much\n",
|
||
"cheaper since we sum over the datapoints in the $k-th$ minibatch and not\n",
|
||
"all $n$ datapoints.\n",
|
||
"\n",
|
||
"## When do we stop?\n",
|
||
"\n",
|
||
"A natural question is when do we stop the search for a new minimum?\n",
|
||
"One possibility is to compute the full gradient after a given number\n",
|
||
"of epochs and check if the norm of the gradient is smaller than some\n",
|
||
"threshold and stop if true. However, the condition that the gradient\n",
|
||
"is zero is valid also for local minima, so this would only tell us\n",
|
||
"that we are close to a local/global minimum. However, we could also\n",
|
||
"evaluate the cost function at this point, store the result and\n",
|
||
"continue the search. If the test kicks in at a later stage we can\n",
|
||
"compare the values of the cost function and keep the $\\beta$ that\n",
|
||
"gave the lowest value.\n",
|
||
"\n",
|
||
"## Slightly different approach\n",
|
||
"\n",
|
||
"Another approach is to let the step length $\\gamma_j$ depend on the\n",
|
||
"number of epochs in such a way that it becomes very small after a\n",
|
||
"reasonable time such that we do not move at all.\n",
|
||
"\n",
|
||
"As an example, let $e = 0,1,2,3,\\cdots$ denote the current epoch and let $t_0, t_1 > 0$ be two fixed numbers. Furthermore, let $t = e \\cdot m + i$ where $m$ is the number of minibatches and $i=0,\\cdots,m-1$. Then the function $$\\gamma_j(t; t_0, t_1) = \\frac{t_0}{t+t_1} $$ goes to zero as the number of epochs gets large. I.e. we start with a step length $\\gamma_j (0; t_0, t_1) = t_0/t_1$ which decays in *time* $t$.\n",
|
||
"\n",
|
||
"In this way we can fix the number of epochs, compute $\\beta$ and\n",
|
||
"evaluate the cost function at the end. Repeating the computation will\n",
|
||
"give a different result since the scheme is random by design. Then we\n",
|
||
"pick the final $\\beta$ that gives the lowest value of the cost\n",
|
||
"function."
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "code",
|
||
"execution_count": 23,
|
||
"metadata": {
|
||
"collapsed": false
|
||
},
|
||
"outputs": [],
|
||
"source": [
|
||
"import numpy as np \n",
|
||
"\n",
|
||
"def step_length(t,t0,t1):\n",
|
||
" return t0/(t+t1)\n",
|
||
"\n",
|
||
"n = 100 #100 datapoints \n",
|
||
"M = 5 #size of each minibatch\n",
|
||
"m = int(n/M) #number of minibatches\n",
|
||
"n_epochs = 500 #number of epochs\n",
|
||
"t0 = 1.0\n",
|
||
"t1 = 10\n",
|
||
"\n",
|
||
"gamma_j = t0/t1\n",
|
||
"j = 0\n",
|
||
"for epoch in range(1,n_epochs+1):\n",
|
||
" for i in range(m):\n",
|
||
" k = np.random.randint(m) #Pick the k-th minibatch at random\n",
|
||
" #Compute the gradient using the data in minibatch Bk\n",
|
||
" #Compute new suggestion for beta\n",
|
||
" t = epoch*m+i\n",
|
||
" gamma_j = step_length(t,t0,t1)\n",
|
||
" j += 1\n",
|
||
"\n",
|
||
"print(\"gamma_j after %d epochs: %g\" % (n_epochs,gamma_j))"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"## Program for stochastic gradient"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "code",
|
||
"execution_count": 24,
|
||
"metadata": {
|
||
"collapsed": false
|
||
},
|
||
"outputs": [],
|
||
"source": [
|
||
"# Importing various packages\n",
|
||
"from math import exp, sqrt\n",
|
||
"from random import random, seed\n",
|
||
"import numpy as np\n",
|
||
"import matplotlib.pyplot as plt\n",
|
||
"from sklearn.linear_model import SGDRegressor\n",
|
||
"\n",
|
||
"x = 2*np.random.rand(100,1)\n",
|
||
"y = 4+3*x+np.random.randn(100,1)\n",
|
||
"\n",
|
||
"xb = np.c_[np.ones((100,1)), x]\n",
|
||
"theta_linreg = np.linalg.inv(xb.T.dot(xb)).dot(xb.T).dot(y)\n",
|
||
"print(\"Own inversion\")\n",
|
||
"print(theta_linreg)\n",
|
||
"sgdreg = SGDRegressor(n_iter = 50, penalty=None, eta0=0.1)\n",
|
||
"sgdreg.fit(x,y.ravel())\n",
|
||
"print(\"sgdreg from scikit\")\n",
|
||
"print(sgdreg.intercept_, sgdreg.coef_)\n",
|
||
"\n",
|
||
"\n",
|
||
"theta = np.random.randn(2,1)\n",
|
||
"\n",
|
||
"eta = 0.1\n",
|
||
"Niterations = 1000\n",
|
||
"m = 100\n",
|
||
"\n",
|
||
"for iter in range(Niterations):\n",
|
||
" gradients = 2.0/m*xb.T.dot(xb.dot(theta)-y)\n",
|
||
" theta -= eta*gradients\n",
|
||
"print(\"theta frm own gd\")\n",
|
||
"print(theta)\n",
|
||
"\n",
|
||
"xnew = np.array([[0],[2]])\n",
|
||
"xbnew = np.c_[np.ones((2,1)), xnew]\n",
|
||
"ypredict = xbnew.dot(theta)\n",
|
||
"ypredict2 = xbnew.dot(theta_linreg)\n",
|
||
"\n",
|
||
"\n",
|
||
"n_epochs = 50\n",
|
||
"t0, t1 = 5, 50\n",
|
||
"m = 100\n",
|
||
"def learning_schedule(t):\n",
|
||
" return t0/(t+t1)\n",
|
||
"\n",
|
||
"theta = np.random.randn(2,1)\n",
|
||
"\n",
|
||
"for epoch in range(n_epochs):\n",
|
||
" for i in range(m):\n",
|
||
" random_index = np.random.randint(m)\n",
|
||
" xi = xb[random_index:random_index+1]\n",
|
||
" yi = y[random_index:random_index+1]\n",
|
||
" gradients = 2 * xi.T.dot(xi.dot(theta)-yi)\n",
|
||
" eta = learning_schedule(epoch*m+i)\n",
|
||
" theta = theta - eta*gradients\n",
|
||
"print(\"theta from own sdg\")\n",
|
||
"print(theta)\n",
|
||
"\n",
|
||
"\n",
|
||
"\n",
|
||
"\n",
|
||
"\n",
|
||
"\n",
|
||
"plt.plot(xnew, ypredict, \"r-\")\n",
|
||
"plt.plot(xnew, ypredict2, \"b-\")\n",
|
||
"plt.plot(x, y ,'ro')\n",
|
||
"plt.axis([0,2.0,0, 15.0])\n",
|
||
"plt.xlabel(r'$x$')\n",
|
||
"plt.ylabel(r'$y$')\n",
|
||
"plt.title(r'Random numbers ')\n",
|
||
"plt.show()"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"## Momentum based methods\n",
|
||
"\n",
|
||
"\n",
|
||
"\n",
|
||
"## Conjugate gradient method\n",
|
||
"In the CG method we define so-called conjugate directions and two vectors \n",
|
||
"$\\hat{s}$ and $\\hat{t}$\n",
|
||
"are said to be\n",
|
||
"conjugate if"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{s}^T\\hat{A}\\hat{t}= 0.\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"The philosophy of the CG method is to perform searches in various conjugate directions\n",
|
||
"of our vectors $\\hat{x}_i$ obeying the above criterion, namely"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{x}_i^T\\hat{A}\\hat{x}_j= 0.\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"Two vectors are conjugate if they are orthogonal with respect to \n",
|
||
"this inner product. Being conjugate is a symmetric relation: if $\\hat{s}$ is conjugate to $\\hat{t}$, then $\\hat{t}$ is conjugate to $\\hat{s}$.\n",
|
||
"\n",
|
||
"\n",
|
||
"\n",
|
||
"## Conjugate gradient method\n",
|
||
"An example is given by the eigenvectors of the matrix"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{v}_i^T\\hat{A}\\hat{v}_j= \\lambda\\hat{v}_i^T\\hat{v}_j,\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"which is zero unless $i=j$.\n",
|
||
"\n",
|
||
"\n",
|
||
"\n",
|
||
"\n",
|
||
"## Conjugate gradient method\n",
|
||
"Assume now that we have a symmetric positive-definite matrix $\\hat{A}$ of size\n",
|
||
"$n\\times n$. At each iteration $i+1$ we obtain the conjugate direction of a vector"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{x}_{i+1}=\\hat{x}_{i}+\\alpha_i\\hat{p}_{i}.\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"We assume that $\\hat{p}_{i}$ is a sequence of $n$ mutually conjugate directions. \n",
|
||
"Then the $\\hat{p}_{i}$ form a basis of $R^n$ and we can expand the solution \n",
|
||
"$ \\hat{A}\\hat{x} = \\hat{b}$ in this basis, namely"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{x} = \\sum^{n}_{i=1} \\alpha_i \\hat{p}_i.\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"## Conjugate gradient method\n",
|
||
"The coefficients are given by"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\mathbf{A}\\mathbf{x} = \\sum^{n}_{i=1} \\alpha_i \\mathbf{A} \\mathbf{p}_i = \\mathbf{b}.\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"Multiplying with $\\hat{p}_k^T$ from the left gives"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{p}_k^T \\hat{A}\\hat{x} = \\sum^{n}_{i=1} \\alpha_i\\hat{p}_k^T \\hat{A}\\hat{p}_i= \\hat{p}_k^T \\hat{b},\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"and we can define the coefficients $\\alpha_k$ as"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\alpha_k = \\frac{\\hat{p}_k^T \\hat{b}}{\\hat{p}_k^T \\hat{A} \\hat{p}_k}\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"## Conjugate gradient method and iterations\n",
|
||
"\n",
|
||
"If we choose the conjugate vectors $\\hat{p}_k$ carefully, \n",
|
||
"then we may not need all of them to obtain a good approximation to the solution \n",
|
||
"$\\hat{x}$. \n",
|
||
"We want to regard the conjugate gradient method as an iterative method. \n",
|
||
"This will us to solve systems where $n$ is so large that the direct \n",
|
||
"method would take too much time.\n",
|
||
"\n",
|
||
"We denote the initial guess for $\\hat{x}$ as $\\hat{x}_0$. \n",
|
||
"We can assume without loss of generality that"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{x}_0=0,\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"or consider the system"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{A}\\hat{z} = \\hat{b}-\\hat{A}\\hat{x}_0,\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"instead.\n",
|
||
"\n",
|
||
"\n",
|
||
"\n",
|
||
"\n",
|
||
"## Conjugate gradient method\n",
|
||
"One can show that the solution $\\hat{x}$ is also the unique minimizer of the quadratic form"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"f(\\hat{x}) = \\frac{1}{2}\\hat{x}^T\\hat{A}\\hat{x} - \\hat{x}^T \\hat{x} , \\quad \\hat{x}\\in\\mathbf{R}^n.\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"This suggests taking the first basis vector $\\hat{p}_1$ \n",
|
||
"to be the gradient of $f$ at $\\hat{x}=\\hat{x}_0$, \n",
|
||
"which equals"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{A}\\hat{x}_0-\\hat{b},\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"and \n",
|
||
"$\\hat{x}_0=0$ it is equal $-\\hat{b}$.\n",
|
||
"The other vectors in the basis will be conjugate to the gradient, \n",
|
||
"hence the name conjugate gradient method.\n",
|
||
"\n",
|
||
"\n",
|
||
"\n",
|
||
"\n",
|
||
"## Conjugate gradient method\n",
|
||
"Let $\\hat{r}_k$ be the residual at the $k$-th step:"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{r}_k=\\hat{b}-\\hat{A}\\hat{x}_k.\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"Note that $\\hat{r}_k$ is the negative gradient of $f$ at \n",
|
||
"$\\hat{x}=\\hat{x}_k$, \n",
|
||
"so the gradient descent method would be to move in the direction $\\hat{r}_k$. \n",
|
||
"Here, we insist that the directions $\\hat{p}_k$ are conjugate to each other, \n",
|
||
"so we take the direction closest to the gradient $\\hat{r}_k$ \n",
|
||
"under the conjugacy constraint. \n",
|
||
"This gives the following expression"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{p}_{k+1}=\\hat{r}_k-\\frac{\\hat{p}_k^T \\hat{A}\\hat{r}_k}{\\hat{p}_k^T\\hat{A}\\hat{p}_k} \\hat{p}_k.\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"## Conjugate gradient method\n",
|
||
"We can also compute the residual iteratively as"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{r}_{k+1}=\\hat{b}-\\hat{A}\\hat{x}_{k+1},\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"which equals"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{b}-\\hat{A}(\\hat{x}_k+\\alpha_k\\hat{p}_k),\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"or"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"(\\hat{b}-\\hat{A}\\hat{x}_k)-\\alpha_k\\hat{A}\\hat{p}_k,\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"which gives"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"\\hat{r}_{k+1}=\\hat{r}_k-\\hat{A}\\hat{p}_{k},\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"## Simple implementation of the Conjugate gradient algorithm"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
" Vector ConjugateGradient(Matrix A, Vector b, Vector x0){\n",
|
||
" int dim = x0.Dimension();\n",
|
||
" const double tolerance = 1.0e-14;\n",
|
||
" Vector x(dim),r(dim),v(dim),z(dim);\n",
|
||
" double c,t,d;\n",
|
||
" \n",
|
||
" x = x0;\n",
|
||
" r = b - A*x;\n",
|
||
" v = r;\n",
|
||
" c = dot(r,r);\n",
|
||
" int i = 0; IterMax = dim;\n",
|
||
" while(i <= IterMax){\n",
|
||
" z = A*v;\n",
|
||
" t = c/dot(v,z);\n",
|
||
" x = x + t*v;\n",
|
||
" r = r - t*z;\n",
|
||
" d = dot(r,r);\n",
|
||
" if(sqrt(d) < tolerance)\n",
|
||
" break;\n",
|
||
" v = r + (d/c)*v;\n",
|
||
" c = d; i++;\n",
|
||
" }\n",
|
||
" return x;\n",
|
||
" } \n"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"## Broyden–Fletcher–Goldfarb–Shanno algorithm\n",
|
||
"The optimization problem is to minimize $f(\\mathbf {x} )$ where $\\mathbf {x}$ is a vector in $R^{n}$, and $f$ is a differentiable scalar function. There are no constraints on the values that $\\mathbf {x}$ can take.\n",
|
||
"\n",
|
||
"The algorithm begins at an initial estimate for the optimal value $\\mathbf {x}_{0}$ and proceeds iteratively to get a better estimate at each stage.\n",
|
||
"\n",
|
||
"The search direction $p_k$ at stage $k$ is given by the solution of the analogue of the Newton equation"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"B_{k}\\mathbf {p} _{k}=-\\nabla f(\\mathbf {x}_{k}),\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"where $B_{k}$ is an approximation to the Hessian matrix, which is\n",
|
||
"updated iteratively at each stage, and $\\nabla f(\\mathbf {x} _{k})$\n",
|
||
"is the gradient of the function\n",
|
||
"evaluated at $x_k$. \n",
|
||
"A line search in the direction $p_k$ is then used to\n",
|
||
"find the next point $x_{k+1}$ by minimising"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"$$\n",
|
||
"f(\\mathbf {x}_{k}+\\alpha \\mathbf {p}_{k}),\n",
|
||
"$$"
|
||
]
|
||
},
|
||
{
|
||
"cell_type": "markdown",
|
||
"metadata": {},
|
||
"source": [
|
||
"over the scalar $\\alpha > 0$."
|
||
]
|
||
}
|
||
],
|
||
"metadata": {},
|
||
"nbformat": 4,
|
||
"nbformat_minor": 2
|
||
}
|