diff --git a/doc/src/week36/exercisesweek36.do.txt b/doc/src/week36/exercisesweek36.do.txt new file mode 100644 index 000000000..9def12dac --- /dev/null +++ b/doc/src/week36/exercisesweek36.do.txt @@ -0,0 +1,191 @@ +TITLE: Exercises week 36 +AUTHOR: September 4-8, 2023 +DATE: Deadline is Sunday September 10 at midnight + + +===== Overarching aims of the exercises this week ===== + +This set of exercises form an important part of the first project. The +analytical exercises deal with the material covered last week on the +mathematical interpretations of ordinary least squares and of Ridge +regression. The numerical exercises can be seen as a continuation of +exercise 3 from week 35, with the inclusion of Ridge regression. This +material enters also the discussions of the first project. + +===== Exercise: Analytical exercises ===== + +The aim here is to derive the expression for the optimal parameters +using Ridge regression. Furthermore, using the singular value +decomposition, we will analyze the difference between the ordinary +least squares approach and Ridge regression. + +The expression for the standard Mean Squared Error (MSE) which we used to define our cost function and the equations for the ordinary least squares (OLS) method, that is +our optimization problem is +!bt +\[ +{\displaystyle \min_{\bm{\beta}\in {\mathbb{R}}^{p}}}\frac{1}{n}\left\{\left(\bm{y}-\bm{X}\bm{\beta}\right)^T\left(\bm{y}-\bm{X}\bm{\beta}\right)\right\}. +\] +!et +or we can state it as +!bt +\[ +{\displaystyle \min_{\bm{\beta}\in +{\mathbb{R}}^{p}}}\frac{1}{n}\sum_{i=0}^{n-1}\left(y_i-\tilde{y}_i\right)^2=\frac{1}{n}\vert\vert \bm{y}-\bm{X}\bm{\beta}\vert\vert_2^2, +\] +!et +where we have used the definition of a norm-2 vector, that is +!bt +\[ +\vert\vert \bm{x}\vert\vert_2 = \sqrt{\sum_i x_i^2}. +\] +!et + +By minimizing the above equation with respect to the parameters +$\bm{\beta}$ we could then obtain an analytical expression for the +parameters $\bm{\beta}$. + +We can add a regularization parameter $\lambda$ by +defining a new cost function to be optimized, that is + +!bt +\[ +{\displaystyle \min_{\bm{\beta}\in +{\mathbb{R}}^{p}}}\frac{1}{n}\vert\vert \bm{y}-\bm{X}\bm{\beta}\vert\vert_2^2+\lambda\vert\vert \bm{\beta}\vert\vert_2^2 +\] +!et + +which leads to the Ridge regression minimization problem where we +require that $\vert\vert \bm{\beta}\vert\vert_2^2\le t$, where $t$ is +a finite number larger than zero. + +=== a) Expression for Ridge regression === +Show that the optimal parameters +!bt +\[ +\hat{\bm{\beta}}_{\mathrm{Ridge}} = \left(\bm{X}^T\bm{X}+\lambda\bm{I}\right)^{-1}\bm{X}^T\bm{y}, +\] +!et + +with $\bm{I}$ being a $p\times p$ identity matrix with the constraint that + +!bt +\[ +\sum_{i=0}^{p-1} \beta_i^2 \leq t, +\] +!et + +with $t$ a finite positive number. + +The ordinary least squares result is +!bt +\[ +\hat{\bm{\beta}}_{\mathrm{OLS}} = \left(\bm{X}^T\bm{X}\right)^{-1}\bm{X}^T\bm{y}, +\] +!et + +=== b) The singular value decomposition === + +Use the singular value decomposition of an $m\times n$ matrix $\bm{X}$ (our design matrix) +!bt +\[ + + +discussed in the lecture notes from week 35 to sh + +We have already analyzed the OLS solutions in terms of the eigenvectors (the columns) of the right singular value matrix $\bm{U}$ as +!bt +\[ +\tilde{\bm{y}}_{\mathrm{OLS}}=\bm{X}\bm{\beta} =\bm{U}\bm{U}^T\bm{y}. +\] +!et + + +For Ridge regression this becomes + +!bt +\[ +\tilde{\bm{y}}_{\mathrm{Ridge}}=\bm{X}\bm{\beta}_{\mathrm{Ridge}} = \bm{U\Sigma V^T}\left(\bm{V}\bm{\Sigma}^2\bm{V}^T+\lambda\bm{I} \right)^{-1}(\bm{U\Sigma V^T})^T\bm{y}=\sum_{j=0}^{p-1}\bm{u}_j\bm{u}_j^T\frac{\sigma_j^2}{\sigma_j^2+\lambda}\bm{y}, +\] +!et + +with the vectors $\bm{u}_j$ being the columns of $\bm{U}$ from the SVD of the matrix $\bm{X}$. + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +===== Exercise: Adding Ridge Regression ===== + + +This exercise is a continuation of exercise 3 from last week. We will use the same function to +generate our data set, still staying with a simple function $y(x)$ +which we want to fit using linear regression, but now extending the +analysis to include the Ridge regression method. + + + +Write your own code for the Ridge method (see chapter 3.4 of Hastie +*et al.*, equations (3.43) and (3.44)) and compute the parametrization +for different values of $\lambda$. Compare and analyze your results +with those from exercise 3. Study the dependence on $\lambda$ while +also varying the strength of the noise in your expression for $y(x)$. + + +Repeat the above but using the functionality of +_Scikit-Learn_. Compare your code with the results from +_Scikit-Learn_. Remember to run with the same random numbers for +generating $x$ and $y$. Observe also that when you compare with _Scikit-Learn_, you need to pay attention to how the intercept is dealt with. + + + +Finally, using _Scikit-Learn_ or your own code, compute also the mean square error, a risk metric corresponding to the expected value of the squared (quadratic) error defined as +!bt +\[ MSE(\hat{y},\hat{\tilde{y}}) = \frac{1}{n} +\sum_{i=0}^{n-1}(y_i-\tilde{y}_i)^2, +\] +!et +and the $R^2$ score function. +If $\tilde{\hat{y}}_i$ is the predicted value of the $i-th$ sample and $y_i$ is the corresponding true value, then the score $R^2$ is defined as +!bt +\[ +R^2(\hat{y}, \tilde{\hat{y}}) = 1 - \frac{\sum_{i=0}^{n - 1} (y_i - \tilde{y}_i)^2}{\sum_{i=0}^{n - 1} (y_i - \bar{y})^2}, +\] +!et +where we have defined the mean value of $\hat{y}$ as +!bt +\[ +\bar{y} = \frac{1}{n} \sum_{i=0}^{n - 1} y_i. +\] +!et +Discuss these quantities as functions of the variable $\lambda$ in Ridge regression. + + + + + + +