cleaning more
This commit is contained in:
@@ -293,7 +293,7 @@ $$
|
||||
\frac{\partial {\cal C}}{\partial \gamma} =-2\sum_{i}\beta x_i(y_i-\beta(1+\gamma x_i))=0.
|
||||
$$
|
||||
|
||||
We can then rewrite these equations as (defining \( \boldsymbol{w}=\boldsymbol{e}+\gamma x_i) \) with \( \boldsymbol{e} \) being the unit vector)
|
||||
We can then rewrite these equations as (defining \( \boldsymbol{w}=\boldsymbol{e}+\gamma \boldsymbol{x}) \) with \( \boldsymbol{e} \) being the unit vector)
|
||||
$$
|
||||
\gamma \boldsymbol{w}^T(\boldsymbol{y}-\beta\gamma \boldsymbol{w})=0,
|
||||
$$
|
||||
@@ -305,7 +305,7 @@ $$
|
||||
|
||||
<p>
|
||||
which leads to \( \gamma =(\boldsymbol{x}^T\boldsymbol{y}-\beta\boldsymbol{x}^T\boldsymbol{e})/(\beta\boldsymbol{x}^T\boldsymbol{x}) \). Inserting
|
||||
for \( \beta \) gives us an equation for \( \gamma \).
|
||||
for \( \beta \) gives us an equation for \( \gamma \). This is a non-linear equation in the unknown \( \gamma \) and has to be solved numerically.
|
||||
|
||||
<p>
|
||||
The solution to these two equations gives us in turn \( \beta_1 \) and \( \gamma_1 \) leading to the new expression for \( f_1(x) \) as
|
||||
|
||||
@@ -269,7 +269,7 @@ MathJax.Hub.Config({
|
||||
<h2 id="___sec55" class="anchor">Gradient Boosting, algorithm </h2>
|
||||
|
||||
<p>
|
||||
Suppose we have a cost function \( C(f)=\sum_{i=0}^{n-1}L(y_i, f(x_i)) \) where \( y_i \) is our target and \( f(x_i) \) the function which is meant to model \( y_i \). The above cost function could be our standard square-error function
|
||||
Suppose we have a cost function \( C(f)=\sum_{i=0}^{n-1}L(y_i, f(x_i)) \) where \( y_i \) is our target and \( f(x_i) \) the function which is meant to model \( y_i \). The above cost function could be our standard squared-error function
|
||||
$$
|
||||
C(\boldsymbol{y},\boldsymbol{f})=\sum_{i=0}^{n-1}(y_i-f(x_i))^2.
|
||||
$$
|
||||
|
||||
@@ -2098,7 +2098,7 @@ $$
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
We can then rewrite these equations as (defining \( \boldsymbol{w}=\boldsymbol{e}+\gamma x_i) \) with \( \boldsymbol{e} \) being the unit vector)
|
||||
We can then rewrite these equations as (defining \( \boldsymbol{w}=\boldsymbol{e}+\gamma \boldsymbol{x}) \) with \( \boldsymbol{e} \) being the unit vector)
|
||||
<p> <br>
|
||||
$$
|
||||
\gamma \boldsymbol{w}^T(\boldsymbol{y}-\beta\gamma \boldsymbol{w})=0,
|
||||
@@ -2114,7 +2114,7 @@ $$
|
||||
|
||||
<p>
|
||||
which leads to \( \gamma =(\boldsymbol{x}^T\boldsymbol{y}-\beta\boldsymbol{x}^T\boldsymbol{e})/(\beta\boldsymbol{x}^T\boldsymbol{x}) \). Inserting
|
||||
for \( \beta \) gives us an equation for \( \gamma \).
|
||||
for \( \beta \) gives us an equation for \( \gamma \). This is a non-linear equation in the unknown \( \gamma \) and has to be solved numerically.
|
||||
|
||||
<p>
|
||||
The solution to these two equations gives us in turn \( \beta_1 \) and \( \gamma_1 \) leading to the new expression for \( f_1(x) \) as
|
||||
@@ -2376,7 +2376,7 @@ See discussion during lecture November 8.
|
||||
<h2 id="___sec55">Gradient Boosting, algorithm </h2>
|
||||
|
||||
<p>
|
||||
Suppose we have a cost function \( C(f)=\sum_{i=0}^{n-1}L(y_i, f(x_i)) \) where \( y_i \) is our target and \( f(x_i) \) the function which is meant to model \( y_i \). The above cost function could be our standard square-error function
|
||||
Suppose we have a cost function \( C(f)=\sum_{i=0}^{n-1}L(y_i, f(x_i)) \) where \( y_i \) is our target and \( f(x_i) \) the function which is meant to model \( y_i \). The above cost function could be our standard squared-error function
|
||||
<p> <br>
|
||||
$$
|
||||
C(\boldsymbol{y},\boldsymbol{f})=\sum_{i=0}^{n-1}(y_i-f(x_i))^2.
|
||||
|
||||
@@ -2094,7 +2094,7 @@ $$
|
||||
\frac{\partial {\cal C}}{\partial \gamma} =-2\sum_{i}\beta x_i(y_i-\beta(1+\gamma x_i))=0.
|
||||
$$
|
||||
|
||||
We can then rewrite these equations as (defining \( \boldsymbol{w}=\boldsymbol{e}+\gamma x_i) \) with \( \boldsymbol{e} \) being the unit vector)
|
||||
We can then rewrite these equations as (defining \( \boldsymbol{w}=\boldsymbol{e}+\gamma \boldsymbol{x}) \) with \( \boldsymbol{e} \) being the unit vector)
|
||||
$$
|
||||
\gamma \boldsymbol{w}^T(\boldsymbol{y}-\beta\gamma \boldsymbol{w})=0,
|
||||
$$
|
||||
@@ -2106,7 +2106,7 @@ $$
|
||||
|
||||
<p>
|
||||
which leads to \( \gamma =(\boldsymbol{x}^T\boldsymbol{y}-\beta\boldsymbol{x}^T\boldsymbol{e})/(\beta\boldsymbol{x}^T\boldsymbol{x}) \). Inserting
|
||||
for \( \beta \) gives us an equation for \( \gamma \).
|
||||
for \( \beta \) gives us an equation for \( \gamma \). This is a non-linear equation in the unknown \( \gamma \) and has to be solved numerically.
|
||||
|
||||
<p>
|
||||
The solution to these two equations gives us in turn \( \beta_1 \) and \( \gamma_1 \) leading to the new expression for \( f_1(x) \) as
|
||||
@@ -2337,7 +2337,7 @@ See discussion during lecture November 8.
|
||||
<h2 id="___sec55">Gradient Boosting, algorithm </h2>
|
||||
|
||||
<p>
|
||||
Suppose we have a cost function \( C(f)=\sum_{i=0}^{n-1}L(y_i, f(x_i)) \) where \( y_i \) is our target and \( f(x_i) \) the function which is meant to model \( y_i \). The above cost function could be our standard square-error function
|
||||
Suppose we have a cost function \( C(f)=\sum_{i=0}^{n-1}L(y_i, f(x_i)) \) where \( y_i \) is our target and \( f(x_i) \) the function which is meant to model \( y_i \). The above cost function could be our standard squared-error function
|
||||
$$
|
||||
C(\boldsymbol{y},\boldsymbol{f})=\sum_{i=0}^{n-1}(y_i-f(x_i))^2.
|
||||
$$
|
||||
|
||||
@@ -2099,7 +2099,7 @@ $$
|
||||
\frac{\partial {\cal C}}{\partial \gamma} =-2\sum_{i}\beta x_i(y_i-\beta(1+\gamma x_i))=0.
|
||||
$$
|
||||
|
||||
We can then rewrite these equations as (defining \( \boldsymbol{w}=\boldsymbol{e}+\gamma x_i) \) with \( \boldsymbol{e} \) being the unit vector)
|
||||
We can then rewrite these equations as (defining \( \boldsymbol{w}=\boldsymbol{e}+\gamma \boldsymbol{x}) \) with \( \boldsymbol{e} \) being the unit vector)
|
||||
$$
|
||||
\gamma \boldsymbol{w}^T(\boldsymbol{y}-\beta\gamma \boldsymbol{w})=0,
|
||||
$$
|
||||
@@ -2111,7 +2111,7 @@ $$
|
||||
|
||||
<p>
|
||||
which leads to \( \gamma =(\boldsymbol{x}^T\boldsymbol{y}-\beta\boldsymbol{x}^T\boldsymbol{e})/(\beta\boldsymbol{x}^T\boldsymbol{x}) \). Inserting
|
||||
for \( \beta \) gives us an equation for \( \gamma \).
|
||||
for \( \beta \) gives us an equation for \( \gamma \). This is a non-linear equation in the unknown \( \gamma \) and has to be solved numerically.
|
||||
|
||||
<p>
|
||||
The solution to these two equations gives us in turn \( \beta_1 \) and \( \gamma_1 \) leading to the new expression for \( f_1(x) \) as
|
||||
@@ -2342,7 +2342,7 @@ See discussion during lecture November 8.
|
||||
<h2 id="___sec55">Gradient Boosting, algorithm </h2>
|
||||
|
||||
<p>
|
||||
Suppose we have a cost function \( C(f)=\sum_{i=0}^{n-1}L(y_i, f(x_i)) \) where \( y_i \) is our target and \( f(x_i) \) the function which is meant to model \( y_i \). The above cost function could be our standard square-error function
|
||||
Suppose we have a cost function \( C(f)=\sum_{i=0}^{n-1}L(y_i, f(x_i)) \) where \( y_i \) is our target and \( f(x_i) \) the function which is meant to model \( y_i \). The above cost function could be our standard squared-error function
|
||||
$$
|
||||
C(\boldsymbol{y},\boldsymbol{f})=\sum_{i=0}^{n-1}(y_i-f(x_i))^2.
|
||||
$$
|
||||
|
||||
@@ -2172,7 +2172,7 @@
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"We can then rewrite these equations as (defining $\\boldsymbol{w}=\\boldsymbol{e}+\\gamma x_i)$ with $\\boldsymbol{e}$ being the unit vector)"
|
||||
"We can then rewrite these equations as (defining $\\boldsymbol{w}=\\boldsymbol{e}+\\gamma \\boldsymbol{x})$ with $\\boldsymbol{e}$ being the unit vector)"
|
||||
]
|
||||
},
|
||||
{
|
||||
@@ -2205,7 +2205,7 @@
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"which leads to $\\gamma =(\\boldsymbol{x}^T\\boldsymbol{y}-\\beta\\boldsymbol{x}^T\\boldsymbol{e})/(\\beta\\boldsymbol{x}^T\\boldsymbol{x})$. Inserting\n",
|
||||
"for $\\beta$ gives us an equation for $\\gamma$.\n",
|
||||
"for $\\beta$ gives us an equation for $\\gamma$. This is a non-linear equation in the unknown $\\gamma$ and has to be solved numerically. \n",
|
||||
"\n",
|
||||
"The solution to these two equations gives us in turn $\\beta_1$ and $\\gamma_1$ leading to the new expression for $f_1(x)$ as\n",
|
||||
"$f_1(x) = \\beta_1(1+\\gamma_1x)$. Doing this $M$ times results in our final estimate for the function $f$. \n",
|
||||
@@ -2569,7 +2569,7 @@
|
||||
"\n",
|
||||
"## Gradient Boosting, algorithm\n",
|
||||
"\n",
|
||||
"Suppose we have a cost function $C(f)=\\sum_{i=0}^{n-1}L(y_i, f(x_i))$ where $y_i$ is our target and $f(x_i)$ the function which is meant to model $y_i$. The above cost function could be our standard square-error function"
|
||||
"Suppose we have a cost function $C(f)=\\sum_{i=0}^{n-1}L(y_i, f(x_i))$ where $y_i$ is our target and $f(x_i)$ the function which is meant to model $y_i$. The above cost function could be our standard squared-error function"
|
||||
]
|
||||
},
|
||||
{
|
||||
|
||||
Binary file not shown.
Binary file not shown.
@@ -1727,7 +1727,7 @@ and
|
||||
\frac{\partial {\cal C}}{\partial \gamma} =-2\sum_{i}\beta x_i(y_i-\beta(1+\gamma x_i))=0.
|
||||
\]
|
||||
!et
|
||||
We can then rewrite these equations as (defining $\bm{w}=\bm{e}+\gamma x_i)$ with $\bm{e}$ being the unit vector)
|
||||
We can then rewrite these equations as (defining $\bm{w}=\bm{e}+\gamma \bm{x})$ with $\bm{e}$ being the unit vector)
|
||||
!bt
|
||||
\[
|
||||
\gamma \bm{w}^T(\bm{y}-\beta\gamma \bm{w})=0,
|
||||
@@ -1741,7 +1741,7 @@ which gives us $\beta = \bm{w}^T\bm{y}/(\bm{w}^T\bm{w})$. Similarly we have
|
||||
!et
|
||||
|
||||
which leads to $\gamma =(\bm{x}^T\bm{y}-\beta\bm{x}^T\bm{e})/(\beta\bm{x}^T\bm{x})$. Inserting
|
||||
for $\beta$ gives us an equation for $\gamma$.
|
||||
for $\beta$ gives us an equation for $\gamma$. This is a non-linear equation in the unknown $\gamma$ and has to be solved numerically.
|
||||
|
||||
The solution to these two equations gives us in turn $\beta_1$ and $\gamma_1$ leading to the new expression for $f_1(x)$ as
|
||||
$f_1(x) = \beta_1(1+\gamma_1x)$. Doing this $M$ times results in our final estimate for the function $f$.
|
||||
@@ -1953,7 +1953,7 @@ See discussion during lecture November 8.
|
||||
!split
|
||||
===== Gradient Boosting, algorithm =====
|
||||
|
||||
Suppose we have a cost function $C(f)=\sum_{i=0}^{n-1}L(y_i, f(x_i))$ where $y_i$ is our target and $f(x_i)$ the function which is meant to model $y_i$. The above cost function could be our standard square-error function
|
||||
Suppose we have a cost function $C(f)=\sum_{i=0}^{n-1}L(y_i, f(x_i))$ where $y_i$ is our target and $f(x_i)$ the function which is meant to model $y_i$. The above cost function could be our standard squared-error function
|
||||
!bt
|
||||
\[
|
||||
C(\bm{y},\bm{f})=\sum_{i=0}^{n-1}(y_i-f(x_i))^2.
|
||||
|
||||
Reference in New Issue
Block a user