Updated regression analysis, sad to come
This commit is contained in:
@@ -75,7 +75,16 @@ div { text-align: justify; text-justify: inter-word; }
|
||||
2,
|
||||
None,
|
||||
'___sec3'),
|
||||
('Optimizing our parameters', 2, None, '___sec4')]}
|
||||
('Optimizing our parameters', 2, None, '___sec4'),
|
||||
('Interpretations and optimizing our parameters',
|
||||
2,
|
||||
None,
|
||||
'___sec5'),
|
||||
('Interpretations and optimizing our parameters',
|
||||
2,
|
||||
None,
|
||||
'___sec6'),
|
||||
('The singular value decompostion', 2, None, '___sec7')]}
|
||||
end of tocinfo -->
|
||||
|
||||
<body>
|
||||
@@ -90,8 +99,8 @@ MathJax.Hub.Config({
|
||||
}
|
||||
});
|
||||
</script>
|
||||
<script type="text/javascript" async
|
||||
src="https://cdnjs.cloudflare.com/ajax/libs/mathjax/2.7.1/MathJax.js?config=TeX-AMS-MML_HTMLorMML">
|
||||
<script type="text/javascript"
|
||||
src="http://cdn.mathjax.org/mathjax/latest/MathJax.js?config=TeX-AMS-MML_HTMLorMML">
|
||||
</script>
|
||||
|
||||
|
||||
@@ -117,7 +126,7 @@ MathJax.Hub.Config({
|
||||
<center>[2] <b>Department of Physics and Astronomy and National Superconducting Cyclotron Laboratory, Michigan State University</b></center>
|
||||
<br>
|
||||
<p>
|
||||
<center><h4>Oct 14, 2017</h4></center> <!-- date -->
|
||||
<center><h4>Oct 17, 2017</h4></center> <!-- date -->
|
||||
<br>
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
@@ -172,19 +181,13 @@ where \( \epsilon_i \) is the error in our approximation.
|
||||
<p>
|
||||
For every set of values \( y_i,x_i \) we have thus the corresponding set of equations
|
||||
$$
|
||||
\begin{align}
|
||||
y_0&=\beta_0+\beta_1x_0^1+\beta_2x_0^2+\dots+\beta_{n-1}x_0^{n-1}+\epsilon_0
|
||||
\label{_auto1}\\
|
||||
y_1&=\beta_0+\beta_1x_1^1+\beta_2x_1^2+\dots+\beta_{n-1}x_1^{n-1}+\epsilon_1
|
||||
\label{_auto2}\\
|
||||
y_2&=\beta_0+\beta_1x_2^1+\beta_2x_2^2+\dots+\beta_{n-1}x_2^{n-1}+\epsilon_2
|
||||
\label{_auto3}\\
|
||||
\dots & \dots
|
||||
\label{_auto4}\\
|
||||
y_{n-1}&=\beta_0+\beta_1x_{n-1}^1+\beta_2x_{n-1}^2+\dots+\beta_1x_{n-1}^{n-1}+\epsilon_{n-1}.
|
||||
\label{_auto5}\\
|
||||
\label{_auto6}
|
||||
\end{align}
|
||||
\begin{align*}
|
||||
y_0&=\beta_0+\beta_1x_0^1+\beta_2x_0^2+\dots+\beta_{n-1}x_0^{n-1}+\epsilon_0\\
|
||||
y_1&=\beta_0+\beta_1x_1^1+\beta_2x_1^2+\dots+\beta_{n-1}x_1^{n-1}+\epsilon_1\\
|
||||
y_2&=\beta_0+\beta_1x_2^1+\beta_2x_2^2+\dots+\beta_{n-1}x_2^{n-1}+\epsilon_2\\
|
||||
\dots & \dots \\
|
||||
y_{n-1}&=\beta_0+\beta_1x_{n-1}^1+\beta_2x_{n-1}^2+\dots+\beta_1x_{n-1}^{n-1}+\epsilon_{n-1}.\\
|
||||
\end{align*}
|
||||
$$
|
||||
|
||||
Defining the vectors
|
||||
@@ -229,23 +232,15 @@ $$
|
||||
We are obviously not limited to the above polynomial. We could replace the various powers of \( x \) with elements of Fourier series, that is, instead of \( x_i^j \) we could have \( \cos{(j x_i)} \) or \( \sin{(j x_i)} \), or time series or other orthogonal functions.
|
||||
For every set of values \( y_i,x_i \) we can then generalize the equations to
|
||||
$$
|
||||
\begin{align}
|
||||
y_0&=\beta_0x_{00}+\beta_1x_{01}+\beta_2x_{02}+\dots+\beta_{n-1}x_{0n-1}+\epsilon_0
|
||||
\label{_auto7}\\
|
||||
y_1&=\beta_0x_{10}+\beta_1x_{11}+\beta_2x_{12}+\dots+\beta_{n-1}x_{1n-1}+\epsilon_1
|
||||
\label{_auto8}\\
|
||||
y_2&=\beta_0x_{20}+\beta_1x_{21}+\beta_2x_{22}+\dots+\beta_{n-1}x_{2n-1}+\epsilon_1
|
||||
\label{_auto9}\\
|
||||
\dots & \dots
|
||||
\label{_auto10}\\
|
||||
y_{i}&=\beta_0x_{i0}+\beta_1x_{i1}+\beta_2x_{i2}+\dots+\beta_{n-1}x_{in-1}+\epsilon_1
|
||||
\label{_auto11}\\
|
||||
\dots & \dots
|
||||
\label{_auto12}\\
|
||||
y_{n-1}&=\beta_0x_{n-1,0}+\beta_1x_{n-1,2}+\beta_2x_{n-1,2}+\dots+\beta_1x_{n-1}^{n-1,n-1}+\epsilon_{n-1}.
|
||||
\label{_auto13}\\
|
||||
\label{_auto14}
|
||||
\end{align}
|
||||
\begin{align*}
|
||||
y_0&=\beta_0x_{00}+\beta_1x_{01}+\beta_2x_{02}+\dots+\beta_{n-1}x_{0n-1}+\epsilon_0\\
|
||||
y_1&=\beta_0x_{10}+\beta_1x_{11}+\beta_2x_{12}+\dots+\beta_{n-1}x_{1n-1}+\epsilon_1\\
|
||||
y_2&=\beta_0x_{20}+\beta_1x_{21}+\beta_2x_{22}+\dots+\beta_{n-1}x_{2n-1}+\epsilon_2\\
|
||||
\dots & \dots \\
|
||||
y_{i}&=\beta_0x_{i0}+\beta_1x_{i1}+\beta_2x_{i2}+\dots+\beta_{n-1}x_{in-1}+\epsilon_i\\
|
||||
\dots & \dots \\
|
||||
y_{n-1}&=\beta_0x_{n-1,0}+\beta_1x_{n-1,2}+\beta_2x_{n-1,2}+\dots+\beta_1x_{n-1}^{n-1,n-1}+\epsilon_{n-1}.\\
|
||||
\end{align*}
|
||||
$$
|
||||
|
||||
We redefine in turn the matrix \( \hat{X} \) as
|
||||
@@ -276,28 +271,131 @@ The left-hand side of this equation forms know. Our error vector \( \hat{\epsilo
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
We have defined the matrix \( \hat{X} \)
|
||||
$$
|
||||
\begin{align}
|
||||
y_0&=\beta_0x_{00}+\beta_1x_{01}+\beta_2x_{02}+\dots+\beta_{n-1}x_{0n-1}+\epsilon_0
|
||||
\label{_auto15}\\
|
||||
y_1&=\beta_0x_{10}+\beta_1x_{11}+\beta_2x_{12}+\dots+\beta_{n-1}x_{1n-1}+\epsilon_1
|
||||
\label{_auto16}\\
|
||||
y_2&=\beta_0x_{20}+\beta_1x_{21}+\beta_2x_{22}+\dots+\beta_{n-1}x_{2n-1}+\epsilon_1
|
||||
\label{_auto17}\\
|
||||
\dots & \dots
|
||||
\label{_auto18}\\
|
||||
y_{i}&=\beta_0x_{i0}+\beta_1x_{i1}+\beta_2x_{i2}+\dots+\beta_{n-1}x_{in-1}+\epsilon_1
|
||||
\label{_auto19}\\
|
||||
\dots & \dots
|
||||
\label{_auto20}\\
|
||||
y_{n-1}&=\beta_0x_{n-1,0}+\beta_1x_{n-1,2}+\beta_2x_{n-1,2}+\dots+\beta_1x_{n-1}^{n-1,n-1}+\epsilon_{n-1}.
|
||||
\label{_auto21}\\
|
||||
\label{_auto22}
|
||||
\end{align}
|
||||
\begin{align*}
|
||||
y_0&=\beta_0x_{00}+\beta_1x_{01}+\beta_2x_{02}+\dots+\beta_{n-1}x_{0n-1}+\epsilon_0\\
|
||||
y_1&=\beta_0x_{10}+\beta_1x_{11}+\beta_2x_{12}+\dots+\beta_{n-1}x_{1n-1}+\epsilon_1\\
|
||||
y_2&=\beta_0x_{20}+\beta_1x_{21}+\beta_2x_{22}+\dots+\beta_{n-1}x_{2n-1}+\epsilon_1\\
|
||||
\dots & \dots \\
|
||||
y_{i}&=\beta_0x_{i0}+\beta_1x_{i1}+\beta_2x_{i2}+\dots+\beta_{n-1}x_{in-1}+\epsilon_1\\
|
||||
\dots & \dots \\
|
||||
y_{n-1}&=\beta_0x_{n-1,0}+\beta_1x_{n-1,2}+\beta_2x_{n-1,2}+\dots+\beta_1x_{n-1,n-1}+\epsilon_{n-1}.\\
|
||||
\end{align*}
|
||||
$$
|
||||
|
||||
We well use this matrix to define the approximation \( \hat{\tilde{y}} \) via the unknown quantity \( \hat{\beta} \) as
|
||||
$$
|
||||
\hat{\tilde{y}}= \hat{X}\hat{\beta},
|
||||
$$
|
||||
|
||||
and in order to find the optimal parameters \( \beta_i \) instead of solving the above linear algebra problem, we define a function which gives a measure of the spread between the values \( y_i \) (which represent hopefully the exact values) and the parametrized values \( \tilde{y}_i \), namely
|
||||
$$
|
||||
Q(\hat{\beta})=\sum_{i=0}^{n-1}\left(y_i-\tilde{y}_i\right)^2=\left(\hat{y}-\hat{\tilde{y}}\right)^T\left(\hat{y}-\hat{\tilde{y}}\right),
|
||||
$$
|
||||
|
||||
or using the matrix \( \hat{X} \) as
|
||||
$$
|
||||
Q(\hat{\beta})=\left(\hat{y}-\hat{X}\hat{\beta}\right)^T\left(\hat{y}-\hat{X}\hat{\beta}\right).
|
||||
$$
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec5">Interpretations and optimizing our parameters </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
The function
|
||||
$$
|
||||
Q(\hat{\beta})=\left(\hat{y}-\hat{X}\hat{\beta}\right)^T\left(\hat{y}-\hat{X}\hat{\beta}\right),
|
||||
$$
|
||||
|
||||
can be linked to the variance of the quantity \( y_i \) if we interpret the latter as the mean value of for example a numerical experiment. When linking below with the maximum likelihood approach below, we will indeed interpret \( y_i \) as a mean value
|
||||
$$
|
||||
y_{i}=\langle y_i \rangle = \beta_0x_{i,0}+\beta_1x_{i,1}+\beta_2x_{i,2}+\dots+\beta_{n-1}x_{i,n-1}+\epsilon_i,
|
||||
$$
|
||||
|
||||
where \( \langle y_i \rangle \) is the mean value. Keep in mind also that till now we have treated \( y_i \) as the exact value. Normally, the response (dependent or outcome) variable \( y_i \) the outcome of a numerical experiment or another type of experiment and is thus only an approximation to the true value. It is then always accompanied by an error estimate, often limited to a statistical error estimate given by the standard deviation discussed earlier. In the discussion here we will treat \( y_i \) as our exact value for the response variable.
|
||||
|
||||
<p>
|
||||
In order to find the parameters \( \beta_i \) we will then minimize the spread of \( Q(\hat{\beta}) \) by requiring
|
||||
$$
|
||||
\frac{\partial Q(\hat{\beta})}{\partial \beta_j} = \frac{\partial }{\partial \beta_j}\left[ \sum_{i=0}^{n-1}\left(y_i-\beta_0x_{i,0}+\beta_1x_{i,1}+\beta_2x_{i,2}+\dots+\beta_{n-1}x_{i,n-1}\right)^2\right]=0,
|
||||
$$
|
||||
|
||||
which results in
|
||||
$$
|
||||
\frac{\partial Q(\hat{\beta})}{\partial \beta_j} = -2\left[ \sum_{i=0}^{n-1}x_{ij}\left(y_i-\beta_0x_{i,0}+\beta_1x_{i,1}+\beta_2x_{i,2}+\dots+\beta_{n-1}x_{i,n-1}\right)\right]=0,
|
||||
$$
|
||||
|
||||
or in a matrix-vector form as
|
||||
$$
|
||||
\frac{\partial Q(\hat{\beta})}{\partial \hat{\beta}} = 0 = \hat{X}^T\left( \hat{y}-\hat{X}\hat{\beta}\right).
|
||||
$$
|
||||
|
||||
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec6">Interpretations and optimizing our parameters </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
We can rewrite
|
||||
$$
|
||||
\frac{\partial Q(\hat{\beta})}{\partial \hat{\beta}} = 0 = \hat{X}^T\left( \hat{y}-\hat{X}\hat{\beta}\right),
|
||||
$$
|
||||
|
||||
as
|
||||
$$
|
||||
\hat{X}^T\hat{y} = \hat{X}^T\hat{X}\hat{\beta},
|
||||
$$
|
||||
|
||||
and if the matrix \( \hat{X}^T\hat{X} \) is invertible we have the solution
|
||||
$$
|
||||
\hat{\beta} =\left(\hat{X}^T\hat{X}\right)^{-1}\hat{X}^T\hat{y}.
|
||||
$$
|
||||
|
||||
The residuals \( \hat{\epsilon} \) are in turn given by
|
||||
$$
|
||||
\hat{\epsilon} = \hat{y}-\hat{\tilde{y}} = \hat{y}-\hat{X}\hat{\beta},
|
||||
$$
|
||||
|
||||
and with
|
||||
$$
|
||||
\hat{X}^T\left( \hat{y}-\hat{X}\hat{\beta}\right)= 0,
|
||||
$$
|
||||
|
||||
we have
|
||||
$$
|
||||
\hat{X}^T\hat{\epsilon}=\hat{X}^T\left( \hat{y}-\hat{X}\hat{\beta}\right)= 0,
|
||||
$$
|
||||
|
||||
meaning that the solution for \( \hat{\beta} \) is the one which minimizes the residuals. Later we will link this with the maximum likelihood approach.
|
||||
|
||||
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
<!-- !split --><br><br><br><br><br><br><br><br><br><br>
|
||||
|
||||
<h2 id="___sec7">The singular value decompostion </h2>
|
||||
<div class="alert alert-block alert-block alert-text-normal">
|
||||
<b></b>
|
||||
<p>
|
||||
Here we derive the equations for the SVD.
|
||||
</div>
|
||||
|
||||
|
||||
<p>
|
||||
|
||||
<!-- ------------------- end of main content --------------- -->
|
||||
|
||||
|
||||
|
||||
Reference in New Issue
Block a user