typos in gradient methods
This commit is contained in:
+225
-223
@@ -26,11 +26,12 @@ direction of the negative gradient $-\nabla F(\mathbf{x})$.
|
||||
It can be shown that if
|
||||
!bt
|
||||
\[
|
||||
\mathbf{x}_{k+1} = \mathbf{x}_k - \gamma_k \nabla F(\mathbf{x}_k), \ \ \gamma_k > 0
|
||||
\mathbf{x}_{k+1} = \mathbf{x}_k - \gamma_k \nabla F(\mathbf{x}_k),
|
||||
\]
|
||||
!et
|
||||
with $\gamma_k > 0$.
|
||||
|
||||
for $\gamma_k$ small enough, then $F(\mathbf{x}_{k+1}) \leq
|
||||
For $\gamma_k$ small enough, then $F(\mathbf{x}_{k+1}) \leq
|
||||
F(\mathbf{x}_k)$. This means that for a sufficiently small $\gamma_k$
|
||||
we are always moving towards smaller function values, i.e a minimum.
|
||||
|
||||
@@ -54,7 +55,7 @@ the learning rate within the context of Machine Learning.
|
||||
!split
|
||||
===== The ideal =====
|
||||
|
||||
Ideally the sequence $\{ \mathbf{x}_k \}_{k=0}$ converges to a global
|
||||
Ideally the sequence $\{\mathbf{x}_k \}_{k=0}$ converges to a global
|
||||
minimum of the function $F$. In general we do not know if we are in a
|
||||
global or local minimum. In the special case when $F$ is a convex
|
||||
function, all local minima are also global minima, so in this case
|
||||
@@ -76,7 +77,8 @@ Note that the gradient is a function of $\mathbf{x} =
|
||||
!split
|
||||
===== The sensitiveness of the gradient descent =====
|
||||
|
||||
GD is sensitive to the choice of learning rate $\gamma_k$. This is due
|
||||
The gradient descent method
|
||||
is sensitive to the choice of learning rate $\gamma_k$. This is due
|
||||
to the fact that we are only guaranteed that $F(\mathbf{x}_{k+1}) \leq
|
||||
F(\mathbf{x}_k)$ for sufficiently small $\gamma_k$. The problem is to
|
||||
determine an optimal learning rate. If the learning rate is chosen too
|
||||
@@ -87,10 +89,217 @@ Many of these shortcomings can be alleviated by introducing
|
||||
randomness. One such method is that of Stochastic Gradient Descent
|
||||
(SGD), see below.
|
||||
|
||||
|
||||
!split
|
||||
===== Convex functions =====
|
||||
|
||||
Ideally we want our cost/loss function to be convex(concave).
|
||||
|
||||
First we give the definition of a convex set: A set $C$ in
|
||||
$\mathbb{R}^n$ is said to be convex if, for all $x$ and $y$ in $C$ and
|
||||
all $t \in (0,1)$ , the point $(1 − t)x + ty$ also belongs to
|
||||
C. Geometrically this means that every point on the line segment
|
||||
connecting $x$ and $y$ is in $C$ as discussed below.
|
||||
|
||||
The convex subsets of $\mathbb{R}$ are the intervals of
|
||||
$\mathbb{R}$. Examples of convex sets of $\mathbb{R}^2$ are the
|
||||
regular polygons (triangles, rectangles, pentagons, etc...).
|
||||
|
||||
!split
|
||||
===== Convex function =====
|
||||
|
||||
_Convex function_: Let $X \subset \mathbb{R}^n$ be a convex set. Assume that the function $f: X \rightarrow \mathbb{R}$ is continuous, then $f$ is said to be convex if $$f(tx_1 + (1-t)x_2) \leq tf(x_1) + (1-t)f(x_2) $$ for all $x_1, x_2 \in X$ and for all $t \in [0,1]$. If $\leq$ is replaced with a strict inequaltiy in the definition, we demand $x_1 \neq x_2$ and $t\in(0,1)$ then $f$ is said to be strictly convex. For a single variable function, convexity means that if you draw a straight line connecting $f(x_1)$ and $f(x_2)$, the value of the function on the interval $[x_1,x_2]$ is always below the line as illustrated below.
|
||||
|
||||
!split
|
||||
===== Conditions on convex functions =====
|
||||
|
||||
In the following we state first and second-order conditions which
|
||||
ensures convexity of a function $f$. We write $D_f$ to denote the
|
||||
domain of $f$, i.e the subset of $R^n$ where $f$ is defined. For more
|
||||
details and proofs we refer to: "S. Boyd and L. Vandenberghe. Convex Optimization. Cambridge University Press":"http://stanford.edu/boyd/cvxbook/, 2004".
|
||||
|
||||
!bblock First order condition
|
||||
Suppose $f$ is differentiable (i.e $\nabla f(x)$ is well defined for
|
||||
all $x$ in the domain of $f$). Then $f$ is convex if and only if $D_f$
|
||||
is a convex set and $$f(y) \geq f(x) + \nabla f(x)^T (y-x) $$ holds
|
||||
for all $x,y \in D_f$. This condition means that for a convex function
|
||||
the first order Taylor expansion (right hand side above) at any point
|
||||
a global under estimator of the function. To convince yourself you can
|
||||
make a drawing of $f(x) = x^2+1$ and draw the tangent line to $f(x)$ and
|
||||
note that it is always below the graph.
|
||||
!eblock
|
||||
|
||||
!bblock Second order condition
|
||||
Assume that $f$ is twice
|
||||
differentiable, i.e the Hessian matrix exists at each point in
|
||||
$D_f$. Then $f$ is convex if and only if $D_f$ is a convex set and its
|
||||
Hessian is positive semi-definite for all $x\in D_f$. For a
|
||||
single-variable function this reduces to $f''(x) \geq 0$. Geometrically this means that $f$ has nonnegative curvature
|
||||
everywhere.
|
||||
!eblock
|
||||
|
||||
This condition is particularly useful since it gives us an procedure for determining if the function under consideration is convex, apart from using the definition.
|
||||
|
||||
!split
|
||||
===== More on convex functions =====
|
||||
|
||||
The next result is of great importance to us and the reason why we are
|
||||
going on about convex functions. In machine learning we frequently
|
||||
have to minimize a loss/cost function in order to find the best
|
||||
parameters for the model we are considering.
|
||||
|
||||
Ideally we want the
|
||||
global minimum (for high-dimensional models it is hard to know
|
||||
if we have local or global minimum). However, if the cost/loss function
|
||||
is convex the following result provides invaluable information:
|
||||
|
||||
!bblock Any minimum is global for convex functions
|
||||
Consider the problem of finding $x \in \mathbb{R}^n$ such that $f(x)$
|
||||
is minimal, where $f$ is convex and differentiable. Then, any point
|
||||
$x^*$ that satisfies $\nabla f(x^*) = 0$ is a global minimum.
|
||||
!eblock
|
||||
|
||||
This result means that if we know that the cost/loss function is convex and we are able to find a minimum, we are guaranteed that it is a global minimum.
|
||||
|
||||
!split
|
||||
===== Some simple problems =====
|
||||
|
||||
o Show that $f(x)=x^2$ is convex for $x \in \mathbb{R}$ using the definition of convexity. Hint: If you re-write the definition, $f$ is convex if the following holds for all $x,y \in D_f$ and any $\lambda \in [0,1]$ $\lambda f(x)+(1-\lambda)f(y)-f(\lambda x + (1-\lambda) y ) \geq 0$.
|
||||
|
||||
o Using the second order condition show that the following functions are convex on the specified domain.
|
||||
* $f(x) = e^x$ is convex for $x \in \mathbb{R}$.
|
||||
* $g(x) = -\ln(x)$ is convex for $x \in (0,\infty)$.
|
||||
o Let $f(x) = x^2$ and $g(x) = e^x$. Show that $f(g(x))$ and $g(f(x))$ is convex for $x \in \mathbb{R}$. Also show that if $f(x)$ is any convex function than $h(x) = e^{f(x)}$ is convex.
|
||||
|
||||
o A norm is any function that satisfy the following properties
|
||||
* $f(\alpha x) = |\alpha| f(x)$ for all $\alpha \in \mathbb{R}$.
|
||||
* $f(x+y) \leq f(x) + f(y)$
|
||||
* $f(x) \leq 0$ for all $x \in \mathbb{R}^n$ with equality if and only if $x = 0$
|
||||
|
||||
Using the definition of convexity, try to show that a function satisfying the properties above is convex (the third condition is not needed to show this).
|
||||
|
||||
!split
|
||||
===== Revisiting our first homework =====
|
||||
|
||||
We will use linear regression as a case study for the gradient descent
|
||||
methods. Linear regression is a great test case for the gradient
|
||||
descent methods discussed in the lectures since it has several
|
||||
desirable properties such as:
|
||||
|
||||
o An analytical solution (recall homework set 1).
|
||||
o The gradient can be computed analytically.
|
||||
o The cost function is convex which guarantees that gradient descent converges for small enough learning rates
|
||||
|
||||
We revisit the example from homework set 1 where we had
|
||||
!bt
|
||||
\[
|
||||
y_i = 5x_i^2 + 0.1\xi_i, \ i=1,\cdots,100
|
||||
\]
|
||||
!et
|
||||
with $x_i \in [0,1] $ chosen randomly with a uniform distribution. Additionally $\xi_i$ represents stochastic noise chosen according to a normal distribution $\cal {N}(0,1)$.
|
||||
The linear regression model is given by
|
||||
!bt
|
||||
\[
|
||||
h_\beta(x) = \hat{y} = \beta_0 + \beta_1 x,
|
||||
\]
|
||||
!et
|
||||
such that
|
||||
!bt
|
||||
\[
|
||||
\hat{y}_i = \beta_0 + \beta_1 x_i.
|
||||
\]
|
||||
!et
|
||||
|
||||
!split
|
||||
===== Gradient descent example =====
|
||||
|
||||
Let $\mathbf{y} = (y_1,\cdots,y_n)^T$, $\mathbf{\hat{y}} = (\hat{y}_1,\cdots,\hat{y}_n)^T$ and $\beta = (\beta_0, \beta_1)^T$
|
||||
|
||||
It is convenient to write $\mathbf{\hat{y}} = X\beta$ where $X \in \mathbb{R}^{100 \times 2} $ is the design matrix given by
|
||||
!bt
|
||||
\[
|
||||
X \equiv \begin{bmatrix}
|
||||
1 & x_1 \\
|
||||
\vdots & \vdots \\
|
||||
1 & x_{100} & \\
|
||||
\end{bmatrix}.
|
||||
\]
|
||||
!et
|
||||
The loss function is given by
|
||||
!bt
|
||||
\[
|
||||
C(\beta) = ||X\beta-\mathbf{y}||^2 = ||X\beta||^2 - 2 \mathbf{y}^T X\beta + ||\mathbf{y}||^2 = \sum_{i=1}^{100} (\beta_0 + \beta_1 x_i)^2 - 2 y_i (\beta_0 + \beta_1 x_i) + y_i^2
|
||||
\]
|
||||
!et
|
||||
and we want to find $\beta$ such that $C(\beta)$ is minimized.
|
||||
|
||||
!split
|
||||
===== The derivative of the cost/loss function =====
|
||||
|
||||
Computing $\partial C(\beta) / \partial \beta_0$ and $\partial C(\beta) / \partial \beta_1$ we can show that the gradient can be written as
|
||||
!bt
|
||||
\[
|
||||
\nabla_{\beta} C(\beta) = (\partial C(\beta) / \partial \beta_0, \partial C(\beta) / \partial \beta_1)^T = 2\begin{bmatrix} \sum_{i=1}^{100} \left(\beta_0+\beta_1x_i-y_i\right) \\
|
||||
\sum_{i=1}^{100}\left( x_i (\beta_0+\beta_1x_i)-y_ix_i\right) \\
|
||||
\end{bmatrix} = 2X^T(X\beta - \mathbf{y}),
|
||||
\]
|
||||
!et
|
||||
where $X$ is the design matrix defined above.
|
||||
|
||||
!split
|
||||
===== The Hessian matrix =====
|
||||
The Hessian matrix of $C(\beta)$ is given by
|
||||
!bt
|
||||
\[
|
||||
\hat{H} \equiv \begin{bmatrix}
|
||||
\frac{\partial^2 C(\beta)}{\partial \beta_0^2} & \frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} \\
|
||||
\frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} & \frac{\partial^2 C(\beta)}{\partial \beta_1^2} & \\
|
||||
\end{bmatrix} = 2X^T X.
|
||||
\]
|
||||
!et
|
||||
This result implies that $C(\beta)$ is a convex function since the matrix $X^T X$ always is positive semi-definite.
|
||||
|
||||
!split
|
||||
===== Simple program =====
|
||||
|
||||
We can now write a program that minimizes $C(\beta)$ using the gradient descent method with a constant learning rate $\gamma$ according to
|
||||
!bt
|
||||
\[
|
||||
\beta_{k+1} = \beta_k - \gamma \nabla_\beta C(\beta_k), \ k=0,1,\cdots
|
||||
\]
|
||||
!et
|
||||
|
||||
We can use the expression we computed for the gradient and let use a
|
||||
$\beta_0$ be chosen randomly and let $\gamma = 0.001$. Stop iterating
|
||||
when $||\nabla_\beta C(\beta_k) || \leq \epsilon = 10^{-8}$.
|
||||
|
||||
And finally we can compare our solution for $\beta$ with the analytic result given by
|
||||
$\beta= (X^TX)^{-1} X^T \mathbf{y}$.
|
||||
!bc pycod
|
||||
import numpy as np
|
||||
|
||||
"""
|
||||
The following setup is just a suggestion, feel free to write it the way you like.
|
||||
"""
|
||||
|
||||
#Setup problem described in the exercise
|
||||
N = 100 #Nr of datapoints
|
||||
M = 2 #Nr of features
|
||||
x = np.random.rand(N) #Uniformly generated x-values in [0,1]
|
||||
y = 5*x**2 + 0.1*np.random.randn(N)
|
||||
X = np.c_[np.ones(N),x] #Construct design matrix
|
||||
|
||||
#Compute beta according to normal equations to compare with GD solution
|
||||
Xt_X_inv = np.linalg.inv(np.dot(X.T,X))
|
||||
Xt_y = np.dot(X.transpose(),y)
|
||||
beta_NE = np.dot(Xt_X_inv,Xt_y)
|
||||
print(beta_NE)
|
||||
!ec
|
||||
|
||||
!split
|
||||
===== Gradient Descent Example =====
|
||||
|
||||
We revisit now our simple linear regression example with a linear polynomial.
|
||||
Another simple example is here
|
||||
!bc pycod
|
||||
|
||||
# Importing various packages
|
||||
@@ -156,214 +365,7 @@ print(sgdreg.intercept_, sgdreg.coef_)
|
||||
|
||||
!ec
|
||||
|
||||
!split
|
||||
===== Convex functions =====
|
||||
|
||||
Ideally we want our cost/loss function to be convex(concave).
|
||||
|
||||
First we give the definition of a convex set: A set $C$ in
|
||||
$\mathbb{R}^n$ is said to be convex if, for all $x$ and $y$ in $C$ and
|
||||
all $t \in (0,1)$ , the point $(1 − t)x + ty$ also belongs to
|
||||
C. Geometrically this means that every point on the line segment
|
||||
connecting $x$ and $y$ is in $C$ as discussed below.
|
||||
|
||||
The convex subsets of $\mathbb{R}$ are the intervals of
|
||||
$\mathbb{R}$. Examples of convex sets of $\mathbb{R}^2$ are the
|
||||
regular polygons (triangles, rectangles, pentagons, etc...).
|
||||
|
||||
!split
|
||||
===== Convex function =====
|
||||
|
||||
_Convex function_: Let $X \subset \mathbb{R}^n$ be a convex set. Assume that the function $f: X \rightarrow \mathbb{R}$ is continuous, then $f$ is said to be convex if $$f(tx_1 + (1-t)x_2) \leq tf(x_1) + (1-t)f(x_2) $$ for all $x_1, x_2 \in X$ and for all $t \in [0,1]$. If $\leq$ is replaced with a strict inequaltiy in the definition, we demand $x_1 \neq x_2$ and $t\in(0,1)$ then $f$ is said to be strictly convex. For a single variable function, convexity means that if you draw a straight line connecting $f(x_1)$ and $f(x_2)$, the value of the function on the interval $[x_1,x_2]$ is always below the line as illustrated below.
|
||||
|
||||
!split
|
||||
===== Conditions on convex functions =====
|
||||
|
||||
In the following we state first and second-order conditions which
|
||||
ensures convexity of a function $f$. We write $D_f$ to denote the
|
||||
domain of $f$, i.e the subset of $R^n$ where $f$ is defined. For more
|
||||
details and proofs we refer to: "S. Boyd and L. Vandenberghe. Convex Optimization. Cambridge University Press":"http://stanford.edu/boyd/cvxbook/, 2004".
|
||||
|
||||
!bblock First order condition
|
||||
Suppose $f$ is differentiable (i.e $\nabla f(x)$ is well defined for
|
||||
all $x$ in the domain of $f$). Then $f$ is convex if and only if $D_f$
|
||||
is a convex set and $$f(y) \geq f(x) + \nabla f(x)^T (y-x) $$ holds
|
||||
for all $x,y \in D_f$. This condition means that for a convex function
|
||||
the first order Taylor expansion (right hand side above) at any point
|
||||
a global under estimator of the function. To convince yourself you can
|
||||
make a drawing of f(x) = x^2+1 and draw the tangent line to $f(x)$ and
|
||||
note that it is always below the graph.
|
||||
!eblock
|
||||
|
||||
!bblock Second order condition
|
||||
Assume that $f$ is twice
|
||||
differentiable, i.e the Hessian matrix exists at each point in
|
||||
$D_f$. Then $f$ is convex if and only if $D_f$ is a convex set and its
|
||||
Hessian is positive semi-definite for all $x\in D_f$. For a
|
||||
single-variable function this reduces to $f''(x) \geq
|
||||
0$. Geometrically this means that $f$ has nonnegative curvature
|
||||
everywhere.
|
||||
!eblock
|
||||
|
||||
This condition is particularly useful since it gives us an procedure for determining if the function under consideration is convex, apart from using the definition.
|
||||
|
||||
!split
|
||||
===== More on convex functions =====
|
||||
|
||||
The next result is of great importance to us and the reason why we are
|
||||
going on about convex functions. In machine learning we frequently
|
||||
have to minimize a loss/cost function in order to find the best
|
||||
parameters for the model we are considering.
|
||||
|
||||
Ideally we want the
|
||||
global minimum (for high-dimensional models it is hard to know
|
||||
if we have local or global minimum). However, if the cost/loss function
|
||||
is convex the following result provides invaluable information:
|
||||
|
||||
!bblock Any minimum is global for convex functions
|
||||
Consider the problem of finding $x \in \mathbb{R}^n$ such that $f(x)$
|
||||
is minimal, where $f$ is convex and differentiable. Then, any point
|
||||
$x^*$ that satisfies $\nabla f(x^*) = 0$ is a global minimum.
|
||||
!eblock
|
||||
|
||||
This result means that if we know that the cost/loss function is convex and we are able to find a minimum, we are guaranteed that it is a global minimum.
|
||||
|
||||
!split
|
||||
===== Some simple problems =====
|
||||
|
||||
o Show that $f(x)=x^2$ is convex for $x \in \mathbb{R}$ using the definition of convexity. Hint: If you re-write the definition, $f$ is convex if the following holds for all $x,y \in D_f$ and any $\lambda \in [0,1] $ $\lambda f(x) + (1-\lambda)f(y) - f(\lambda x + (1-\lambda) y ) \geq 0. $
|
||||
|
||||
o Using the second order condition show that the following functions are convex on the specified domain.
|
||||
* $f(x) = e^x$ is convex for $x \in \mathbb{R}$.
|
||||
* $g(x) = -\ln(x)$ is convex for $x \in (0,\infty)$.
|
||||
o Let $f(x) = x^2$ and $g(x) = e^x$. Show that $f(g(x))$ and $g(f(x))$ is convex for $x \in \mathbb{R}$. Also show that if $f(x)$ is any convex function than $h(x) = e^{f(x)}$ is convex.
|
||||
|
||||
o A norm is any function that satisfy the following properties
|
||||
* $f(\alpha x) = |\alpha| f(x)$ for all $\alpha \in \mathbb{R}$.
|
||||
* $f(x+y) \leq f(x) + f(y)$
|
||||
* $f(x) \leq 0$ for all $x \in \mathbb{R}^n$ with equality if and only if $x = 0$
|
||||
|
||||
Using the definition of convexity, try to show that a function satisfying the properties above is convex (the third condition is not needed to show this).
|
||||
|
||||
!split
|
||||
===== Revisiting our first homework =====
|
||||
|
||||
We will use linear regression as a case study for the gradient descent
|
||||
methods. Linear regression is a great test case for the gradient
|
||||
descent methods discussed in the lectures since it has several
|
||||
desirable properties such as:
|
||||
|
||||
o An analytical solution (recall homework set 1).
|
||||
o The gradient can be computed analytically.
|
||||
o The cost function is convex which guarantees that gradient descent converges for small enough learning rates
|
||||
|
||||
We revisit the example from homework set 1 where we had
|
||||
!bt
|
||||
\[
|
||||
y_i = 5x_i^2 + 0.1\xi_i, \ i=1,\cdots,100
|
||||
\]
|
||||
!et
|
||||
with $x_i \in [0,1] $ chosen randomly with a uniform distribution. Additionally $\xi_i$ represents stochastic noise chosen according to a normal distribution $\cal {N}(0,1)$.
|
||||
The linear regression model is given by
|
||||
!bt
|
||||
\[
|
||||
h_\beta(x) = \hat{y} = \beta_0 + \beta_1 x,
|
||||
\]
|
||||
!et
|
||||
such that
|
||||
!bt
|
||||
\[
|
||||
\hat{y}_i = \beta_0 + \beta_1 x_i.
|
||||
\]
|
||||
!et
|
||||
|
||||
!split
|
||||
===== Gradient descent example =====
|
||||
|
||||
Let $\mathbf{y} = (y_1,\cdots,y_n)^T$, $\mathbf{\hat{y}} = (\hat{y}_1,\cdots,\hat{y}_n)^T$ and $\beta = (\beta_0, \beta_1)^T$
|
||||
|
||||
t is convenient to write $\mathbf{\hat{y}} = X\beta$ where $X \in \mathbb{R}^{100 \times 2} $ is the design matrix given by
|
||||
!bt
|
||||
\[
|
||||
\begin{equation}
|
||||
X \equiv \begin{bmatrix}
|
||||
1 & x_1 \\
|
||||
\vdots & \vdots \\
|
||||
1 & x_{100} & \\
|
||||
\end{bmatrix}.
|
||||
\end{equation}
|
||||
\]
|
||||
!et
|
||||
The loss function is given by
|
||||
!bt
|
||||
\[
|
||||
C(\beta) = ||X\beta-\mathbf{y}||^2 = ||X\beta||^2 - 2 \mathbf{y}^T X\beta + ||\mathbf{y}||^2 = \sum_{i=1}^{100} (\beta_0 + \beta_1 x_i)^2 - 2 y_i (\beta_0 + \beta_1 x_i) + y_i^2
|
||||
\]
|
||||
!et
|
||||
and we want to find $\beta$ such that $C(\beta)$ is minimized.
|
||||
|
||||
!split
|
||||
===== The derivative of the cost/loss function =====
|
||||
|
||||
Computing $\partial C(\beta) / \partial \beta_0$ and $\partial C(\beta) / \partial \beta_1$ we can show that the gradient can be written as
|
||||
!bt
|
||||
\[
|
||||
\nabla_\beta C(\beta) = (\partial C(\beta) / \partial \beta_0, \partial C(\beta) / \partial \beta_1)^T = 2\begin{bmatrix} \sum_{i=1}^{100} \left(\beta_0+\beta_1x_i-y_i\right) \\
|
||||
\sum_{i=1}^{100}\left( x_i (\beta_0+\beta_1x_i)-y_ix_i\right) \\
|
||||
\end{bmatrix} = 2X^T(X\beta - \mathbf{y}),
|
||||
\]
|
||||
!et
|
||||
where $X$ is the design matrix defined above.
|
||||
|
||||
!split
|
||||
===== The Hessian matrix =====
|
||||
The Hessian matrix of $C(\beta)$ is given by
|
||||
!bt
|
||||
\[
|
||||
\hat{H} \equiv \begin{bmatrix}
|
||||
\frac{\partial^2 C(\beta)}{\partial \beta_0^2} & \frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} \\
|
||||
\frac{\partial^2 C(\beta)}{\partial \beta_0 \partial \beta_1} & \frac{\partial^2 C(\beta)}{\partial \beta_1^2} & \\
|
||||
\end{bmatrix} = 2X^T X.
|
||||
\]
|
||||
!et
|
||||
This result implies that $C(\beta)$ is a convex function since the matrix $X^T X$ always is positive semi-definite.
|
||||
|
||||
!split
|
||||
===== Simple program =====
|
||||
|
||||
We can now write a program that minimizes $C(\beta)$ using the gradient descent method with a constant learning rate $\gamma$ according to
|
||||
!bt
|
||||
\[
|
||||
\beta_{k+1} = \beta_k - \gamma \nabla_\beta C(\beta_k), \ k=0,1,\cdots
|
||||
\]
|
||||
!et
|
||||
|
||||
We can use the expression we computed for the gradient and let use a
|
||||
$\beta_0$ be chosen randomly and let $\gamma = 0.001$. Stop iterating
|
||||
when $||\nabla_\beta C(\beta_k) || < \epsilon = 10^{-8}$.
|
||||
|
||||
And finally we can compare our solution for $\beta$ with the analytic result given by
|
||||
$\beta= (X^TX)^{-1} X^T \mathbf{y}$.
|
||||
!bc pycod
|
||||
import numpy as np
|
||||
|
||||
"""
|
||||
The following setup is just a suggestion, feel free to write it the way you like.
|
||||
"""
|
||||
|
||||
#Setup problem described in the exercise
|
||||
N = 100 #Nr of datapoints
|
||||
M = 2 #Nr of features
|
||||
x = np.random.rand(N) #Uniformly generated x-values in [0,1]
|
||||
y = 5*x**2 + 0.1*np.random.randn(N)
|
||||
X = np.c_[np.ones(N),x] #Construct design matrix
|
||||
|
||||
#Compute beta according to normal equations to compare with GD solution
|
||||
Xt_X_inv = np.linalg.inv(np.dot(X.T,X))
|
||||
Xt_y = np.dot(X.transpose(),y)
|
||||
beta_NE = np.dot(Xt_X_inv,Xt_y)
|
||||
print(beta_NE)
|
||||
!ec
|
||||
|
||||
!split
|
||||
===== Gradient descent and Ridge =====
|
||||
@@ -432,7 +434,7 @@ the shortcomings of the Gradient descent method discussed above.
|
||||
|
||||
The underlying idea of SGD comes from the observation that the cost
|
||||
function, which we want to minimize, can almost always be written as a
|
||||
sum over $n$ datapoints $\{\mathbf{x}_i\}_{i=1}^n$,
|
||||
sum over $n$ data points $\{\mathbf{x}_i\}_{i=1}^n$,
|
||||
!bt
|
||||
\[
|
||||
C(\mathbf{\beta}) = \sum_{i=1}^n c_i(\mathbf{x}_i,
|
||||
@@ -454,27 +456,27 @@ computed as a sum over $i$-gradients
|
||||
|
||||
Stochasticity/randomness is introduced by only taking the
|
||||
gradient on a subset of the data called minibatches. If there are $n$
|
||||
datapoints and the size of each minibatch is $M$, there will be $n/M$
|
||||
data points and the size of each minibatch is $M$, there will be $n/M$
|
||||
minibatches. We denote these minibatches by $B_k$ where
|
||||
$k=1,\cdots,n/M$.
|
||||
|
||||
!split
|
||||
===== SGD example =====
|
||||
As an example, suppose we have $10$ datapoints $( \mathbf{x}_1,
|
||||
\cdots, \mathbf{x}_{10} )$ and we choose to have $M=5$ minibathces,
|
||||
then each minibatch contains two datapoints. In particular we have
|
||||
As an example, suppose we have $10$ data points $(\mathbf{x}_1,\cdots, \mathbf{x}_{10})$
|
||||
and we choose to have $M=5$ minibathces,
|
||||
then each minibatch contains two data points. In particular we have
|
||||
$B_1 = (\mathbf{x}_1,\mathbf{x}_2), \cdots, B_5 =
|
||||
(\mathbf{x}_9,\mathbf{x}_{10})$. Note that if you choose $M=1$ you
|
||||
have only a single batch with all datapoints and on the other extreme,
|
||||
have only a single batch with all data points and on the other extreme,
|
||||
you may choose $M=n$ resulting in a minibatch for each datapoint, i.e
|
||||
$B_k = \mathbf{x}_k$.
|
||||
|
||||
The idea is now to approximate the gradient by replacing the sum over
|
||||
all datapoints with a sum over the datapoints in one the minibatches
|
||||
all data points with a sum over the data points in one the minibatches
|
||||
picked at random in each gradient descent step
|
||||
!bt
|
||||
\[
|
||||
\nabla_\beta
|
||||
\nabla_{\beta}
|
||||
C(\mathbf{\beta}) = \sum_{i=1}^n \nabla_\beta c_i(\mathbf{x}_i,
|
||||
\mathbf{\beta}) \rightarrow \sum_{i \in B_k}^n \nabla_\beta
|
||||
c_i(\mathbf{x}_i, \mathbf{\beta}).
|
||||
@@ -522,8 +524,8 @@ Taking the gradient only on a subset of the data has two important
|
||||
benefits. First, it introduces randomness which decreases the chance
|
||||
that our opmization scheme gets stuck in a local minima. Second, if
|
||||
the size of the minibatches are small relative to the number of
|
||||
datapoints ($M < n$), the computation of the gradient is much
|
||||
cheaper since we sum over the datapoints in the k-th minibatch and not
|
||||
datapoints ($M < n$), the computation of the gradient is much
|
||||
cheaper since we sum over the datapoints in the $k-th$ minibatch and not
|
||||
all $n$ datapoints.
|
||||
|
||||
!split
|
||||
@@ -547,7 +549,7 @@ Another approach is to let the step length $\gamma_j$ depend on the
|
||||
number of epochs in such a way that it becomes very small after a
|
||||
reasonable time such that we do not move at all.
|
||||
|
||||
As an example, let $e = 0,1,2,3,\cdots$ denote the current epoch and let $t_0, t_1 > 0$ be two fixed numbers. Furthermore, let $t = e \cdot m + i$ where $m$ is the number of minibatches and $i=0,\cdots,m-1$. Then the function $$\gamma_j(t; t_0, t_1) = \frac{t_0}{t+t_1} $$ goes to zero as the number of epochs gets large. I.e. we start with a step length $\gamma_j (0; t_0, t_1) = t_0/t_1$ which decays in *time* $t$.
|
||||
As an example, let $e = 0,1,2,3,\cdots$ denote the current epoch and let $t_0, t_1 > 0$ be two fixed numbers. Furthermore, let $t = e \cdot m + i$ where $m$ is the number of minibatches and $i=0,\cdots,m-1$. Then the function $$\gamma_j(t; t_0, t_1) = \frac{t_0}{t+t_1} $$ goes to zero as the number of epochs gets large. I.e. we start with a step length $\gamma_j (0; t_0, t_1) = t_0/t_1$ which decays in *time* $t$.
|
||||
|
||||
In this way we can fix the number of epochs, compute $\beta$ and
|
||||
evaluate the cost function at the end. Repeating the computation will
|
||||
|
||||
Reference in New Issue
Block a user