update week 35
This commit is contained in:
@@ -159,6 +159,8 @@ Automatically generated HTML file from DocOnce source
|
||||
2,
|
||||
None,
|
||||
'mathematical-interpretation-of-ordinary-least-squares'),
|
||||
('Residual Error', 2, None, 'residual-error'),
|
||||
('Simple case', 2, None, 'simple-case'),
|
||||
('The singular value decomposition',
|
||||
2,
|
||||
None,
|
||||
@@ -305,32 +307,34 @@ MathJax.Hub.Config({
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs033.html#testing-the-means-squared-error-as-function-of-complexity" style="font-size: 80%;"><b>Testing the Means Squared Error as function of Complexity</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs034.html#more-preprocessing-examples-franke-function-and-regression" style="font-size: 80%;"><b>More preprocessing examples, Franke function and regression</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs035.html#mathematical-interpretation-of-ordinary-least-squares" style="font-size: 80%;"><b>Mathematical Interpretation of Ordinary Least Squares</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs036.html#the-singular-value-decomposition" style="font-size: 80%;"><b>The singular value decomposition</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs037.html#linear-regression-problems" style="font-size: 80%;"><b>Linear Regression Problems</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs038.html#fixing-the-singularity" style="font-size: 80%;"><b>Fixing the singularity</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs039.html#basic-math-of-the-svd" style="font-size: 80%;"><b>Basic math of the SVD</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs040.html#the-svd-a-fantastic-algorithm" style="font-size: 80%;"><b>The SVD, a Fantastic Algorithm</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs041.html#economy-size-svd" style="font-size: 80%;"><b>Economy-size SVD</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs042.html#codes-for-the-svd" style="font-size: 80%;"><b>Codes for the SVD</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs043.html#mathematical-properties" style="font-size: 80%;"><b>Mathematical Properties</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs044.html#friday-september-3" style="font-size: 80%;"><b>Friday September 3</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs045.html#ridge-and-lasso-regression" style="font-size: 80%;"><b>Ridge and LASSO Regression</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs046.html#more-on-ridge-regression" style="font-size: 80%;"><b>More on Ridge Regression</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs047.html#interpreting-the-ridge-results" style="font-size: 80%;"><b>Interpreting the Ridge results</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs048.html#more-interpretations" style="font-size: 80%;"><b>More interpretations</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs049.html#a-better-understanding-of-regularization" style="font-size: 80%;"><b>A better understanding of regularization</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs050.html#decomposing-the-ols-and-ridge-expressions" style="font-size: 80%;"><b>Decomposing the OLS and Ridge expressions</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs051.html#introducing-the-covariance-and-correlation-functions" style="font-size: 80%;"><b>Introducing the Covariance and Correlation functions</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs052.html#correlation-function-and-design-feature-matrix" style="font-size: 80%;"><b>Correlation Function and Design/Feature Matrix</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs053.html#covariance-matrix-examples" style="font-size: 80%;"><b>Covariance Matrix Examples</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs054.html#correlation-matrix" style="font-size: 80%;"><b>Correlation Matrix</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs055.html#correlation-matrix-with-pandas" style="font-size: 80%;"><b>Correlation Matrix with Pandas</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs056.html#correlation-matrix-with-pandas-and-the-franke-function" style="font-size: 80%;"><b>Correlation Matrix with Pandas and the Franke function</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs057.html#rewriting-the-covariance-and-or-correlation-matrix" style="font-size: 80%;"><b>Rewriting the Covariance and/or Correlation Matrix</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs058.html#linking-with-svd" style="font-size: 80%;"><b>Linking with SVD</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs059.html#exercises-for-week-37-september-6-10" style="font-size: 80%;"><b>Exercises for week 37, September 6-10</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs059.html#exercise-1-adding-ridge-and-lasso-regression" style="font-size: 80%;"><b>Exercise 1: Adding Ridge and Lasso Regression</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs059.html#exercise-linear-regression-for-a-two-dimensional-function" style="font-size: 80%;"> Exercise: Linear Regression for a two-dimensional function</a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs036.html#residual-error" style="font-size: 80%;"><b>Residual Error</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs037.html#simple-case" style="font-size: 80%;"><b>Simple case</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs038.html#the-singular-value-decomposition" style="font-size: 80%;"><b>The singular value decomposition</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs039.html#linear-regression-problems" style="font-size: 80%;"><b>Linear Regression Problems</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs040.html#fixing-the-singularity" style="font-size: 80%;"><b>Fixing the singularity</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs041.html#basic-math-of-the-svd" style="font-size: 80%;"><b>Basic math of the SVD</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs042.html#the-svd-a-fantastic-algorithm" style="font-size: 80%;"><b>The SVD, a Fantastic Algorithm</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs043.html#economy-size-svd" style="font-size: 80%;"><b>Economy-size SVD</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs044.html#codes-for-the-svd" style="font-size: 80%;"><b>Codes for the SVD</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs045.html#mathematical-properties" style="font-size: 80%;"><b>Mathematical Properties</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs046.html#friday-september-3" style="font-size: 80%;"><b>Friday September 3</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs047.html#ridge-and-lasso-regression" style="font-size: 80%;"><b>Ridge and LASSO Regression</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs048.html#more-on-ridge-regression" style="font-size: 80%;"><b>More on Ridge Regression</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs049.html#interpreting-the-ridge-results" style="font-size: 80%;"><b>Interpreting the Ridge results</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs050.html#more-interpretations" style="font-size: 80%;"><b>More interpretations</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs051.html#a-better-understanding-of-regularization" style="font-size: 80%;"><b>A better understanding of regularization</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs052.html#decomposing-the-ols-and-ridge-expressions" style="font-size: 80%;"><b>Decomposing the OLS and Ridge expressions</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs053.html#introducing-the-covariance-and-correlation-functions" style="font-size: 80%;"><b>Introducing the Covariance and Correlation functions</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs054.html#correlation-function-and-design-feature-matrix" style="font-size: 80%;"><b>Correlation Function and Design/Feature Matrix</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs055.html#covariance-matrix-examples" style="font-size: 80%;"><b>Covariance Matrix Examples</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs056.html#correlation-matrix" style="font-size: 80%;"><b>Correlation Matrix</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs057.html#correlation-matrix-with-pandas" style="font-size: 80%;"><b>Correlation Matrix with Pandas</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs058.html#correlation-matrix-with-pandas-and-the-franke-function" style="font-size: 80%;"><b>Correlation Matrix with Pandas and the Franke function</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs059.html#rewriting-the-covariance-and-or-correlation-matrix" style="font-size: 80%;"><b>Rewriting the Covariance and/or Correlation Matrix</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs060.html#linking-with-svd" style="font-size: 80%;"><b>Linking with SVD</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs061.html#exercises-for-week-37-september-6-10" style="font-size: 80%;"><b>Exercises for week 37, September 6-10</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs061.html#exercise-1-adding-ridge-and-lasso-regression" style="font-size: 80%;"><b>Exercise 1: Adding Ridge and Lasso Regression</b></a></li>
|
||||
<!-- navigation toc: --> <li><a href="._week35-bs061.html#exercise-linear-regression-for-a-two-dimensional-function" style="font-size: 80%;"> Exercise: Linear Regression for a two-dimensional function</a></li>
|
||||
|
||||
</ul>
|
||||
</li>
|
||||
@@ -389,7 +393,7 @@ MathJax.Hub.Config({
|
||||
<li><a href="._week35-bs008.html">9</a></li>
|
||||
<li><a href="._week35-bs009.html">10</a></li>
|
||||
<li><a href="">...</a></li>
|
||||
<li><a href="._week35-bs059.html">60</a></li>
|
||||
<li><a href="._week35-bs061.html">62</a></li>
|
||||
<li><a href="._week35-bs001.html">»</a></li>
|
||||
</ul>
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
@@ -178,7 +178,7 @@ The main topics on Thursday are:
|
||||
<p><li> Repetition from last week on linear regression</li>
|
||||
<p><li> Discussion of how to prepare data and examples of applications of linear regression</li>
|
||||
<p><li> Mathematical interpretations of Linear Regression</li>
|
||||
<p><li> Start discussing Ridge regression and Singular Value Decomposition</li>
|
||||
<p><li> Start discussing Ridge and Lasso regression and Singular Value Decomposition</li>
|
||||
</ol>
|
||||
</section>
|
||||
|
||||
@@ -1522,6 +1522,89 @@ clf = skl.LinearRegression().fit(X_train_scaled, y_train)
|
||||
|
||||
<p>
|
||||
What is presented here is a mathematical analysis of various regression algorithms (ordinary least squares, Ridge and Lasso Regression). The analysis is based on an important algorithm in linear algebra, the so-called Singular Value Decomposition (SVD).
|
||||
|
||||
<p>
|
||||
We have shown that in ordinary least squares that the optimal parameter \( \beta \) are given by
|
||||
|
||||
<p> <br>
|
||||
$$
|
||||
\hat{\boldsymbol{\beta}} = \left(\boldsymbol{X}^T\boldsymbol{X}\right^{-1}\boldsymbol{X}^T\boldsymbol{y}.
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
This means that our best model is defined as
|
||||
|
||||
<p> <br>
|
||||
$$
|
||||
\tilde{\boldsymbol{y}}=\boldsymbol{X}\hat{\boldsymbol{\beta}} = \boldsymbol{X}\left(\boldsymbol{X}^T\boldsymbol{X}\right^{-1}\boldsymbol{X}^T\boldsymbol{y}.
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
We define now a matrix
|
||||
<p> <br>
|
||||
$$
|
||||
\boldsymbol{A}=\boldsymbol{X}\left(\boldsymbol{X}^T\boldsymbol{X}\right^{-1}\boldsymbol{X}^T.
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
It means that we can rewrite
|
||||
<p> <br>
|
||||
$$
|
||||
\tilde{\boldsymbol{y}}=\boldsymbol{X}\hat{\boldsymbol{\beta}} = \boldsymbol{A}\boldsymbol{y}.
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
The matrix \( \boldsymbol{A} \) has the important property that \( \boldsymbol{A}^2=\boldsymbol{A} \). This is the definition of a projection matrix.
|
||||
We can then interpret that our optimal model \( \tilde{\boldsymbol{y}} \) ir represented by an orthogonal (it has to be a square matrix like ours, if not we have an oblique projection matrix ) projection of \( \boldsymbol{y} \) onto a space defined by the column vectors of \( \boldsymbol{X} \).
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="residual-error">Residual Error </h2>
|
||||
|
||||
<p>
|
||||
We have defined the residual error as
|
||||
<p> <br>
|
||||
$$
|
||||
\boldsymbol{\epsilon}=\boldsymbol{y}-\tilde{\boldsymbol{y}}=(\boldsymbol{I}-\boldsymbol{X}\left(\boldsymbol{X}^T\boldsymbol{X}\right^{-1}\boldsymbol{X}^T)\boldsymbol{y}.
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
The residual errors are then the projections of \( \boldsymbol{y} \) onto the orthogonal components of the space defined by the column vectors of \( \boldsymbol{X} \).
|
||||
</section>
|
||||
|
||||
|
||||
<section>
|
||||
<h2 id="simple-case">Simple case </h2>
|
||||
|
||||
<p>
|
||||
If the matrix \( \boldsymbol{X} \) is an orthogonal (or unitary in case of complex values) matrix, we have
|
||||
|
||||
<p> <br>
|
||||
$$
|
||||
\boldsymbol{X}^T\boldsymbol{X}=\boldsymbol{X}\boldsymbol{X}^T = \boldsymbol{I}.
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
<p>
|
||||
In this case the matrix \( \boldsymbol{A} \) becomes
|
||||
<p> <br>
|
||||
$$
|
||||
\boldsymbol{A}=\boldsymbol{X}\left(\boldsymbol{X}^T\boldsymbol{X}\right^{-1}\boldsymbol{X}^T)=\boldsymbol{I},
|
||||
$$
|
||||
<p> <br>
|
||||
|
||||
and we have the obvious case
|
||||
<p> <br>
|
||||
$$
|
||||
\boldsymbol{\epsilon}=\boldsymbol{y}-\tilde{\boldsymbol{y}}=0.
|
||||
$$
|
||||
<p> <br>
|
||||
</section>
|
||||
|
||||
|
||||
|
||||
@@ -179,6 +179,8 @@ div { text-align: justify; text-justify: inter-word; }
|
||||
2,
|
||||
None,
|
||||
'mathematical-interpretation-of-ordinary-least-squares'),
|
||||
('Residual Error', 2, None, 'residual-error'),
|
||||
('Simple case', 2, None, 'simple-case'),
|
||||
('The singular value decomposition',
|
||||
2,
|
||||
None,
|
||||
@@ -317,7 +319,7 @@ The main topics on Thursday are:
|
||||
<li> Repetition from last week on linear regression</li>
|
||||
<li> Discussion of how to prepare data and examples of applications of linear regression</li>
|
||||
<li> Mathematical interpretations of Linear Regression</li>
|
||||
<li> Start discussing Ridge regression and Singular Value Decomposition</li>
|
||||
<li> Start discussing Ridge and Lasso regression and Singular Value Decomposition</li>
|
||||
</ol>
|
||||
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
@@ -1614,6 +1616,73 @@ clf = skl.LinearRegression().fit(X_train_scaled, y_train)
|
||||
<p>
|
||||
What is presented here is a mathematical analysis of various regression algorithms (ordinary least squares, Ridge and Lasso Regression). The analysis is based on an important algorithm in linear algebra, the so-called Singular Value Decomposition (SVD).
|
||||
|
||||
<p>
|
||||
We have shown that in ordinary least squares that the optimal parameter \( \beta \) are given by
|
||||
|
||||
$$
|
||||
\hat{\boldsymbol{\beta}} = \left(\boldsymbol{X}^T\boldsymbol{X}\right^{-1}\boldsymbol{X}^T\boldsymbol{y}.
|
||||
$$
|
||||
|
||||
<p>
|
||||
This means that our best model is defined as
|
||||
|
||||
$$
|
||||
\tilde{\boldsymbol{y}}=\boldsymbol{X}\hat{\boldsymbol{\beta}} = \boldsymbol{X}\left(\boldsymbol{X}^T\boldsymbol{X}\right^{-1}\boldsymbol{X}^T\boldsymbol{y}.
|
||||
$$
|
||||
|
||||
<p>
|
||||
We define now a matrix
|
||||
$$
|
||||
\boldsymbol{A}=\boldsymbol{X}\left(\boldsymbol{X}^T\boldsymbol{X}\right^{-1}\boldsymbol{X}^T.
|
||||
$$
|
||||
|
||||
<p>
|
||||
It means that we can rewrite
|
||||
$$
|
||||
\tilde{\boldsymbol{y}}=\boldsymbol{X}\hat{\boldsymbol{\beta}} = \boldsymbol{A}\boldsymbol{y}.
|
||||
$$
|
||||
|
||||
<p>
|
||||
The matrix \( \boldsymbol{A} \) has the important property that \( \boldsymbol{A}^2=\boldsymbol{A} \). This is the definition of a projection matrix.
|
||||
We can then interpret that our optimal model \( \tilde{\boldsymbol{y}} \) ir represented by an orthogonal (it has to be a square matrix like ours, if not we have an oblique projection matrix ) projection of \( \boldsymbol{y} \) onto a space defined by the column vectors of \( \boldsymbol{X} \).
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="residual-error">Residual Error </h2>
|
||||
|
||||
<p>
|
||||
We have defined the residual error as
|
||||
$$
|
||||
\boldsymbol{\epsilon}=\boldsymbol{y}-\tilde{\boldsymbol{y}}=(\boldsymbol{I}-\boldsymbol{X}\left(\boldsymbol{X}^T\boldsymbol{X}\right^{-1}\boldsymbol{X}^T)\boldsymbol{y}.
|
||||
$$
|
||||
|
||||
<p>
|
||||
The residual errors are then the projections of \( \boldsymbol{y} \) onto the orthogonal components of the space defined by the column vectors of \( \boldsymbol{X} \).
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="simple-case">Simple case </h2>
|
||||
|
||||
<p>
|
||||
If the matrix \( \boldsymbol{X} \) is an orthogonal (or unitary in case of complex values) matrix, we have
|
||||
|
||||
$$
|
||||
\boldsymbol{X}^T\boldsymbol{X}=\boldsymbol{X}\boldsymbol{X}^T = \boldsymbol{I}.
|
||||
$$
|
||||
|
||||
<p>
|
||||
In this case the matrix \( \boldsymbol{A} \) becomes
|
||||
$$
|
||||
\boldsymbol{A}=\boldsymbol{X}\left(\boldsymbol{X}^T\boldsymbol{X}\right^{-1}\boldsymbol{X}^T)=\boldsymbol{I},
|
||||
$$
|
||||
|
||||
and we have the obvious case
|
||||
$$
|
||||
\boldsymbol{\epsilon}=\boldsymbol{y}-\tilde{\boldsymbol{y}}=0.
|
||||
$$
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
|
||||
@@ -184,6 +184,8 @@ div { text-align: justify; text-justify: inter-word; }
|
||||
2,
|
||||
None,
|
||||
'mathematical-interpretation-of-ordinary-least-squares'),
|
||||
('Residual Error', 2, None, 'residual-error'),
|
||||
('Simple case', 2, None, 'simple-case'),
|
||||
('The singular value decomposition',
|
||||
2,
|
||||
None,
|
||||
@@ -322,7 +324,7 @@ The main topics on Thursday are:
|
||||
<li> Repetition from last week on linear regression</li>
|
||||
<li> Discussion of how to prepare data and examples of applications of linear regression</li>
|
||||
<li> Mathematical interpretations of Linear Regression</li>
|
||||
<li> Start discussing Ridge regression and Singular Value Decomposition</li>
|
||||
<li> Start discussing Ridge and Lasso regression and Singular Value Decomposition</li>
|
||||
</ol>
|
||||
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
@@ -1619,6 +1621,73 @@ clf <span style="color: #666666">=</span> skl<span style="color: #666666">.</spa
|
||||
<p>
|
||||
What is presented here is a mathematical analysis of various regression algorithms (ordinary least squares, Ridge and Lasso Regression). The analysis is based on an important algorithm in linear algebra, the so-called Singular Value Decomposition (SVD).
|
||||
|
||||
<p>
|
||||
We have shown that in ordinary least squares that the optimal parameter \( \beta \) are given by
|
||||
|
||||
$$
|
||||
\hat{\boldsymbol{\beta}} = \left(\boldsymbol{X}^T\boldsymbol{X}\right^{-1}\boldsymbol{X}^T\boldsymbol{y}.
|
||||
$$
|
||||
|
||||
<p>
|
||||
This means that our best model is defined as
|
||||
|
||||
$$
|
||||
\tilde{\boldsymbol{y}}=\boldsymbol{X}\hat{\boldsymbol{\beta}} = \boldsymbol{X}\left(\boldsymbol{X}^T\boldsymbol{X}\right^{-1}\boldsymbol{X}^T\boldsymbol{y}.
|
||||
$$
|
||||
|
||||
<p>
|
||||
We define now a matrix
|
||||
$$
|
||||
\boldsymbol{A}=\boldsymbol{X}\left(\boldsymbol{X}^T\boldsymbol{X}\right^{-1}\boldsymbol{X}^T.
|
||||
$$
|
||||
|
||||
<p>
|
||||
It means that we can rewrite
|
||||
$$
|
||||
\tilde{\boldsymbol{y}}=\boldsymbol{X}\hat{\boldsymbol{\beta}} = \boldsymbol{A}\boldsymbol{y}.
|
||||
$$
|
||||
|
||||
<p>
|
||||
The matrix \( \boldsymbol{A} \) has the important property that \( \boldsymbol{A}^2=\boldsymbol{A} \). This is the definition of a projection matrix.
|
||||
We can then interpret that our optimal model \( \tilde{\boldsymbol{y}} \) ir represented by an orthogonal (it has to be a square matrix like ours, if not we have an oblique projection matrix ) projection of \( \boldsymbol{y} \) onto a space defined by the column vectors of \( \boldsymbol{X} \).
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="residual-error">Residual Error </h2>
|
||||
|
||||
<p>
|
||||
We have defined the residual error as
|
||||
$$
|
||||
\boldsymbol{\epsilon}=\boldsymbol{y}-\tilde{\boldsymbol{y}}=(\boldsymbol{I}-\boldsymbol{X}\left(\boldsymbol{X}^T\boldsymbol{X}\right^{-1}\boldsymbol{X}^T)\boldsymbol{y}.
|
||||
$$
|
||||
|
||||
<p>
|
||||
The residual errors are then the projections of \( \boldsymbol{y} \) onto the orthogonal components of the space defined by the column vectors of \( \boldsymbol{X} \).
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="simple-case">Simple case </h2>
|
||||
|
||||
<p>
|
||||
If the matrix \( \boldsymbol{X} \) is an orthogonal (or unitary in case of complex values) matrix, we have
|
||||
|
||||
$$
|
||||
\boldsymbol{X}^T\boldsymbol{X}=\boldsymbol{X}\boldsymbol{X}^T = \boldsymbol{I}.
|
||||
$$
|
||||
|
||||
<p>
|
||||
In this case the matrix \( \boldsymbol{A} \) becomes
|
||||
$$
|
||||
\boldsymbol{A}=\boldsymbol{X}\left(\boldsymbol{X}^T\boldsymbol{X}\right^{-1}\boldsymbol{X}^T)=\boldsymbol{I},
|
||||
$$
|
||||
|
||||
and we have the obvious case
|
||||
$$
|
||||
\boldsymbol{\epsilon}=\boldsymbol{y}-\tilde{\boldsymbol{y}}=0.
|
||||
$$
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
|
||||
Binary file not shown.
@@ -32,7 +32,7 @@
|
||||
"\n",
|
||||
"3. Mathematical interpretations of Linear Regression\n",
|
||||
"\n",
|
||||
"4. Start discussing Ridge regression and Singular Value Decomposition\n",
|
||||
"4. Start discussing Ridge and Lasso regression and Singular Value Decomposition\n",
|
||||
"\n",
|
||||
"## Why Linear Regression (aka Ordinary Least Squares and family), repeat from last week\n",
|
||||
"\n",
|
||||
@@ -1858,6 +1858,143 @@
|
||||
"What is presented here is a mathematical analysis of various regression algorithms (ordinary least squares, Ridge and Lasso Regression). The analysis is based on an important algorithm in linear algebra, the so-called Singular Value Decomposition (SVD). \n",
|
||||
"\n",
|
||||
"\n",
|
||||
"We have shown that in ordinary least squares that the optimal parameter $\\beta$ are given by"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\hat{\\boldsymbol{\\beta}} = \\left(\\boldsymbol{X}^T\\boldsymbol{X}\\right^{-1}\\boldsymbol{X}^T\\boldsymbol{y}.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"This means that our best model is defined as"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\tilde{\\boldsymbol{y}}=\\boldsymbol{X}\\hat{\\boldsymbol{\\beta}} = \\boldsymbol{X}\\left(\\boldsymbol{X}^T\\boldsymbol{X}\\right^{-1}\\boldsymbol{X}^T\\boldsymbol{y}.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"We define now a matrix"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\boldsymbol{A}=\\boldsymbol{X}\\left(\\boldsymbol{X}^T\\boldsymbol{X}\\right^{-1}\\boldsymbol{X}^T.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"It means that we can rewrite"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\tilde{\\boldsymbol{y}}=\\boldsymbol{X}\\hat{\\boldsymbol{\\beta}} = \\boldsymbol{A}\\boldsymbol{y}.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"The matrix $\\boldsymbol{A}$ has the important property that $\\boldsymbol{A}^2=\\boldsymbol{A}$. This is the definition of a projection matrix.\n",
|
||||
"We can then interpret that our optimal model $\\tilde{\\boldsymbol{y}}$ ir represented by an orthogonal (it has to be a square matrix like ours, if not we have an oblique projection matrix ) projection of $\\boldsymbol{y}$ onto a space defined by the column vectors of $\\boldsymbol{X}$.\n",
|
||||
"\n",
|
||||
"## Residual Error\n",
|
||||
"\n",
|
||||
"We have defined the residual error as"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\boldsymbol{\\epsilon}=\\boldsymbol{y}-\\tilde{\\boldsymbol{y}}=(\\boldsymbol{I}-\\boldsymbol{X}\\left(\\boldsymbol{X}^T\\boldsymbol{X}\\right^{-1}\\boldsymbol{X}^T)\\boldsymbol{y}.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"The residual errors are then the projections of $\\boldsymbol{y}$ onto the orthogonal components of the space defined by the column vectors of $\\boldsymbol{X}$.\n",
|
||||
"\n",
|
||||
"## Simple case\n",
|
||||
"\n",
|
||||
"If the matrix $\\boldsymbol{X}$ is an orthogonal (or unitary in case of complex values) matrix, we have"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\boldsymbol{X}^T\\boldsymbol{X}=\\boldsymbol{X}\\boldsymbol{X}^T = \\boldsymbol{I}.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"In this case the matrix $\\boldsymbol{A}$ becomes"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\boldsymbol{A}=\\boldsymbol{X}\\left(\\boldsymbol{X}^T\\boldsymbol{X}\\right^{-1}\\boldsymbol{X}^T)=\\boldsymbol{I},\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"and we have the obvious case"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"$$\n",
|
||||
"\\boldsymbol{\\epsilon}=\\boldsymbol{y}-\\tilde{\\boldsymbol{y}}=0.\n",
|
||||
"$$"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## The singular value decomposition\n",
|
||||
"\n",
|
||||
"\n",
|
||||
|
||||
@@ -17,7 +17,7 @@ The main topics on Thursday are:
|
||||
o Repetition from last week on linear regression
|
||||
o Discussion of how to prepare data and examples of applications of linear regression
|
||||
o Mathematical interpretations of Linear Regression
|
||||
o Start discussing Ridge regression and Singular Value Decomposition
|
||||
o Start discussing Ridge and Lasso regression and Singular Value Decomposition
|
||||
|
||||
!split
|
||||
===== Why Linear Regression (aka Ordinary Least Squares and family), repeat from last week =====
|
||||
@@ -1157,6 +1157,77 @@ print("R2 score for scaled data: {:.2f}".format(clf.score(X_test_scaled,y_test)
|
||||
What is presented here is a mathematical analysis of various regression algorithms (ordinary least squares, Ridge and Lasso Regression). The analysis is based on an important algorithm in linear algebra, the so-called Singular Value Decomposition (SVD).
|
||||
|
||||
|
||||
We have shown that in ordinary least squares that the optimal parameter $\beta$ are given by
|
||||
|
||||
!bt
|
||||
\[
|
||||
\hat{\bm{\beta}} = \left(\bm{X}^T\bm{X}\right^{-1}\bm{X}^T\bm{y}.
|
||||
\]
|
||||
!et
|
||||
|
||||
This means that our best model is defined as
|
||||
|
||||
!bt
|
||||
\[
|
||||
\tilde{\bm{y}}=\bm{X}\hat{\bm{\beta}} = \bm{X}\left(\bm{X}^T\bm{X}\right^{-1}\bm{X}^T\bm{y}.
|
||||
\]
|
||||
!et
|
||||
|
||||
We define now a matrix
|
||||
!bt
|
||||
\[
|
||||
\bm{A}=\bm{X}\left(\bm{X}^T\bm{X}\right^{-1}\bm{X}^T.
|
||||
\]
|
||||
!et
|
||||
|
||||
It means that we can rewrite
|
||||
!bt
|
||||
\[
|
||||
\tilde{\bm{y}}=\bm{X}\hat{\bm{\beta}} = \bm{A}\bm{y}.
|
||||
\]
|
||||
!et
|
||||
|
||||
The matrix $\bm{A}$ has the important property that $\bm{A}^2=\bm{A}$. This is the definition of a projection matrix.
|
||||
We can then interpret that our optimal model $\tilde{\bm{y}}$ ir represented by an orthogonal (it has to be a square matrix like ours, if not we have an oblique projection matrix ) projection of $\bm{y}$ onto a space defined by the column vectors of $\bm{X}$.
|
||||
|
||||
!split
|
||||
===== Residual Error =====
|
||||
|
||||
We have defined the residual error as
|
||||
!bt
|
||||
\[
|
||||
\bm{\epsilon}=\bm{y}-\tilde{\bm{y}}=(\bm{I}-\bm{X}\left(\bm{X}^T\bm{X}\right^{-1}\bm{X}^T)\bm{y}.
|
||||
\]
|
||||
!et
|
||||
|
||||
The residual errors are then the projections of $\bm{y}$ onto the orthogonal components of the space defined by the column vectors of $\bm{X}$.
|
||||
|
||||
!split
|
||||
===== Simple case =====
|
||||
|
||||
If the matrix $\bm{X}$ is an orthogonal (or unitary in case of complex values) matrix, we have
|
||||
|
||||
!bt
|
||||
\[
|
||||
\bm{X}^T\bm{X}=\bm{X}\bm{X}^T = \bm{I}.
|
||||
\]
|
||||
!et
|
||||
|
||||
In this case the matrix $\bm{A}$ becomes
|
||||
!bt
|
||||
\[
|
||||
\bm{A}=\bm{X}\left(\bm{X}^T\bm{X}\right^{-1}\bm{X}^T)=\bm{I},
|
||||
\]
|
||||
!et
|
||||
and we have the obvious case
|
||||
!bt
|
||||
\[
|
||||
\bm{\epsilon}=\bm{y}-\tilde{\bm{y}}=0.
|
||||
\]
|
||||
!et
|
||||
|
||||
|
||||
|
||||
!split
|
||||
===== The singular value decomposition =====
|
||||
|
||||
|
||||
Reference in New Issue
Block a user