Finish project. Hopefully
This commit is contained in:
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
|
Before Width: | Height: | Size: 111 KiB After Width: | Height: | Size: 121 KiB |
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
@@ -0,0 +1 @@
|
|||||||
|
In summary it was shown that in the context of big data and complex models to describe relations in the data there are many challenges to overcome. The choice of model, on what metric to optimize the model -- i.e. the cost function -- and how to perform the optimization are all crucial choices that can have a significant impact on the performance of the model. The use of regularization techniques can help to prevent overfitting and improve the generalization performance of the model. The number of degrees of freedom in the model directly influcences the bias-variance trade-off and only with a good balance between the two can a model perform well on unseen data. Finally, the choice of optimization algorithm can have a significant impact on the computational cost of the models training process. Using good optimization algorithms can help to speed up the convergence of the models parameters. Combining this with other innovative techniques such as stochastic gradient descent can help to further reduce the computational cost of training complex models on large datasets.
|
||||||
@@ -1,6 +1,6 @@
|
|||||||
Optimizing numerical models to model a set of data has been a challenge of the natural sciences for many decades. But with the introduction of artificial neural networks in all disciplines over the last years a special focus on those optimization techniques is relevant.
|
Optimizing numerical models to model a set of data has been a challenge of the natural sciences for many decades. But with the introduction of artificial neural networks in all disciplines over the last years a special focus on those optimization techniques is relevant.
|
||||||
|
|
||||||
In this paper different methods of such optimization are studied in the context of large datasets and big data with possibly noise-affected input data. To render the best possible fits a multitude of cost functions, minimization algorithms and other approaches to reduce the computation associated with the optimization process are studied and evaluated.
|
In this paper different methods of such optimization are studied in the context of large datasets and big data with possibly noise-affected input data. To render the best possible fits a multitude of cost functions, minimization algorithms and other approaches to reduce the computation associated with the optimization process are studied and evaluated. Reducing the computational cost is also important due improve the economical and ecological footprint of training large models on big data, as the spread of datadriven methods in all areas of life is increasing rapidly.
|
||||||
|
|
||||||
In the following section the theoretical background of the methods used in this paper is presented. In \cref{sec:methods} secondly the implementation of these methods is discussed. The results of applying these methods to different datasets are presented in \cref{sec:results}. Finally, a conclusion is drawn in \cref{sec:conclusion} and an outlook on possible future work is given in \cref{sec:perspective}.
|
In the following section the theoretical background of the methods used in this paper is presented. In \cref{sec:methods} secondly the implementation of these methods is discussed. The results of applying these methods to different datasets are presented in \cref{sec:results}. Finally, a conclusion is drawn in \cref{sec:conclusion} and an outlook on possible future work is given in \cref{sec:perspective}.
|
||||||
|
|
||||||
|
|||||||
+82
-23
@@ -15,25 +15,25 @@ To find the optimal solution to a linear model in the form
|
|||||||
\begin{equation} \label{eq:linear_model}
|
\begin{equation} \label{eq:linear_model}
|
||||||
X \vec \theta = \vec y, \quad X \in \mathbb{R}^{n\times p}, \vec y \in \mathbb{R}^n, \vec \theta \in \mathbb{R}^p,
|
X \vec \theta = \vec y, \quad X \in \mathbb{R}^{n\times p}, \vec y \in \mathbb{R}^n, \vec \theta \in \mathbb{R}^p,
|
||||||
\end{equation}
|
\end{equation}
|
||||||
the ordinary least squares (OLS) method can be used. The OLS method minimizes the squared deviation between the model prediction and the actual data. This is done by minimizing the cost function
|
the ordinary least squares (OLS) method can be used. The OLS method minimizes the squared deviation between the model prediction and the actual data\cite{hastie_elements_2009}. This is done by minimizing the cost function
|
||||||
\begin{equation} \label{eq:ols_cost}
|
\begin{equation} \label{eq:ols_cost}
|
||||||
\cols(\vec \theta) = \frac{1}{n} \norm{X\vec \theta - \vec y}_2^2.
|
\cols(\vec \theta) = \frac{1}{n} \norm{X\vec \theta - \vec y}_2^2.
|
||||||
\end{equation}
|
\end{equation}
|
||||||
This method has been widely used in the past ultimately also because it has an analytical solution in the case that the matrix product $X^TX$ is invertible, i.e. $\det(X^TX) \neq 0$. In this case the optimal parameters $\vec \theta_{OLS}$ can be found by solving the normal equations
|
This method has been widely used in the past ultimately also because it has an analytical solution in the case that the matrix product $X^TX$ is invertible, i.e. $\det(X^TX) \neq 0$. In this case the optimal parameters $\vec \theta_{OLS}$ can be found by solving the normal equations \cite{hastie_elements_2009,elstner_lecture_2025}
|
||||||
\begin{equation} \label{eq:ols_solution}
|
\begin{equation} \label{eq:ols_solution}
|
||||||
\vec \theta_{OLS} = (X^TX)^{-1}X^T \vec y.
|
\vec \theta_{OLS} = (X^TX)^{-1}X^T \vec y.
|
||||||
\end{equation}
|
\end{equation}
|
||||||
Taking the derivative of the cost function in \cref{eq:ols_cost} with respect to the parameters $\vec \theta$ and inserting the solution in \cref{eq:ols_solution} shows that this is indeed a minimum of the cost function: \todo{Expand on the derivation here?}
|
Taking the derivative of the cost function in \cref{eq:ols_cost} with respect to the parameters $\vec \theta$ and inserting the solution in \cref{eq:ols_solution} shows that this is indeed a minimum of the cost function:
|
||||||
\begin{equation}
|
\begin{equation}
|
||||||
\frac{\partial \cols}{\partial \vec \theta}\Big\rvert_{\vec \theta_{OLS}} = \frac{2}{n} X^T(X\vec \theta - \vec y)\Big\rvert_{\vec \theta_{OLS}} = 0.
|
\frac{\partial \cols}{\partial \vec \theta}\Big\rvert_{\vec \theta_{OLS}} = \frac{2}{n} X^T(X\vec \theta - \vec y)\Big\rvert_{\vec \theta_{OLS}} = 0.
|
||||||
\end{equation}
|
\end{equation}
|
||||||
|
|
||||||
\subsubsection{Bias-Variance Tradeoff and Regularization}
|
\subsubsection{Bias-Variance Tradeoff and Regularization} \label{sec:bias_variance_tradeoff}
|
||||||
The cost function of the ordinary least squares method can also be identified as the mean squared error (MSE) between the model prediction and the actual data
|
The cost function of the ordinary least squares method can also be identified as the mean squared error (MSE) between the model prediction and the actual data
|
||||||
\begin{equation} \label{eq:mse}
|
\begin{equation} \label{eq:mse}
|
||||||
\mse(\vec y, \vec{\tilde y}) = \frac{1}{n} \sum_{i=0}^{n-1} (y_i - \tilde y_i)^2 = \E \left[(y - \tilde y)^2\right],
|
\mse(\vec y, \vec{\tilde y}) = \frac{1}{n} \sum_{i=0}^{n-1} (y_i - \tilde y_i)^2 = \E \left[(y - \tilde y)^2\right],
|
||||||
\end{equation}
|
\end{equation}
|
||||||
where $\vec{\tilde y} = X \vec \theta$ is the model prediction. $\E$ denotes the expectation value. The MSE can be decomposed into three components, the squared bias, the variance and the irreducible error $\sigma^2$ as
|
where $\vec{\tilde y} = X \vec \theta$ is the model prediction. $\E$ denotes the expectation value. The MSE can be decomposed into three components, the squared bias, the variance and the irreducible error $\sigma^2$ as \cite{bishop_pattern_2006,hastie_elements_2009}
|
||||||
\begin{equation} \label{eq:bias_variance_decomp}
|
\begin{equation} \label{eq:bias_variance_decomp}
|
||||||
\mse(\vec y, \vec{\tilde y}) = \bias^2(\vec y, \vec{\tilde y}) + \var(\vec{\tilde y}) + \sigma^2,
|
\mse(\vec y, \vec{\tilde y}) = \bias^2(\vec y, \vec{\tilde y}) + \var(\vec{\tilde y}) + \sigma^2,
|
||||||
\end{equation}
|
\end{equation}
|
||||||
@@ -45,7 +45,7 @@ and the variance as
|
|||||||
\begin{equation} \label{eq:variance}
|
\begin{equation} \label{eq:variance}
|
||||||
\var(\vec{\tilde y}) = \E\left[(\vec{\tilde y} - \E[\vec{\tilde y}])^2\right].
|
\var(\vec{\tilde y}) = \E\left[(\vec{\tilde y} - \E[\vec{\tilde y}])^2\right].
|
||||||
\end{equation}
|
\end{equation}
|
||||||
The irreducible error $\sigma^2$ is the variance of the noise in the data and cannot be reduced by any model. The bias-variance decomposition in \cref{eq:bias_variance_decomp} shows that there is a tradeoff between bias and variance when trying to minimize the MSE. A model with low complexity will have a high bias but low variance, while a model with high complexity will have a low bias but high variance. This is known as the bias-variance tradeoff. The OLS method can lead to overfitting of the data, especially when the number of parameters $p$ is large compared to the number of data points $n$. In this case, the model will fit the noise in the data rather than the underlying trend, leading to a low bias but high variance. On the contrary, a model with too few parameters will not be able to capture the underlying trend in the data, leading to a high bias but low variance. To mitigate the effects of overfitting, regularization techniques such as Ridge and Lasso regression can be used. These techniques add a penalty term to the cost function that penalizes large values of the parameters $\vec \theta$. The cost function for Ridge regression is given by
|
The irreducible error $\sigma^2$ is the variance of the noise in the data and cannot be reduced by any model. The bias-variance decomposition in \cref{eq:bias_variance_decomp} shows that there is a tradeoff between bias and variance when trying to minimize the MSE. A model with low complexity will have a high bias but low variance, while a model with high complexity will have a low bias but high variance. This is known as the bias-variance tradeoff. The OLS method can lead to overfitting of the data, especially when the number of parameters $p$ is large compared to the number of data points $n$. In this case, the model will fit the noise in the data rather than the underlying trend, leading to a low bias but high variance. On the contrary, a model with too few parameters will not be able to capture the underlying trend in the data, leading to a high bias but low variance\cite{elstner_lecture_2025,hastie_elements_2009}. To mitigate the effects of overfitting, regularization techniques such as Ridge and Lasso regression can be used. These techniques add a penalty term to the cost function that penalizes large values of the parameters $\vec \theta$\cite{lekhansh_lasso_2024}. The cost function for Ridge regression is given by
|
||||||
\begin{equation} \label{eq:ridge_cost}
|
\begin{equation} \label{eq:ridge_cost}
|
||||||
\colr(\vec \theta) = \frac{1}{n} \norm{X\vec \theta - \vec y}_2^2 + \lambda \norm{\vec \theta}_2^2,
|
\colr(\vec \theta) = \frac{1}{n} \norm{X\vec \theta - \vec y}_2^2 + \lambda \norm{\vec \theta}_2^2,
|
||||||
\end{equation}
|
\end{equation}
|
||||||
@@ -53,28 +53,38 @@ where $\lambda$ is the regularization parameter that controls the strength of th
|
|||||||
\begin{equation} \label{eq:lasso_cost}
|
\begin{equation} \label{eq:lasso_cost}
|
||||||
\coll(\vec \theta) = \frac{1}{n} \norm{X\vec \theta - \vec y}_2^2 + \lambda \norm{\vec \theta}_1,
|
\coll(\vec \theta) = \frac{1}{n} \norm{X\vec \theta - \vec y}_2^2 + \lambda \norm{\vec \theta}_1,
|
||||||
\end{equation}
|
\end{equation}
|
||||||
where the $L_1$ norm is defined as $\norm{\vec \theta}_1 = \sum_{i=0}^{p-1} |\theta_i|$. Both Ridge and Lasso regression can help to reduce the variance of the model by penalizing large values of the parameters $\vec \theta$, leading to a different bias-variance tradeoff. For Ridge regression, we can derive the effective number of parameters also called degrees of freedom as
|
where the $L_1$ norm is defined as $\norm{\vec \theta}_1 = \sum_{i=0}^{p-1} |\theta_i|$\cite{lekhansh_lasso_2024,elstner_lecture_2025}. Both Ridge and Lasso regression can help to reduce the variance of the model by penalizing large values of the parameters $\vec \theta$, leading to a different bias-variance tradeoff. For Ridge regression, we can derive the effective number of parameters also called degrees of freedom as \cite{hastie_elements_2009}
|
||||||
\begin{equation} \label{eq:ridge_dof}
|
\begin{equation} \label{eq:ridge_dof}
|
||||||
n_\text{effective} = \sum_{i=1}^{n_\text{features}} \frac{d_i^2}{d_i^2 + \lambda},
|
n_\text{effective} = \sum_{i=1}^{n_\text{features}} \frac{d_i^2}{d_i^2 + \lambda},
|
||||||
\end{equation}
|
\end{equation}
|
||||||
where $d_i$ are the singular values of the design matrix $X$. Thus regularization effectively reduces the number of parameters in the model, leading to a bias-variance tradeoff that can be tuned by the regularization parameter $\lambda$. While OLS models can be analytically optimized using \cref{eq:ols_solution} for the case where $X^TX$ is invertible, the same is true for Ridge regression in the case where $X^TX + \lambda I$ is invertible. The optimal parameters for Ridge regression can be found by solving the modified normal equations
|
where $d_i$ are the singular values of the design matrix $X$. Thus regularization effectively reduces the number of parameters in the model, leading to a bias-variance tradeoff that can be tuned by the regularization parameter $\lambda$. While OLS models can be analytically optimized using \cref{eq:ols_solution} for the case where $X^TX$ is invertible, the same is true for Ridge regression in the case where $X^TX + \lambda I$ is invertible. The optimal parameters for Ridge regression can be found by solving the modified normal equations\cite{elstner_lecture_2025}
|
||||||
\begin{equation} \label{eq:ridge_solution}
|
\begin{equation} \label{eq:ridge_solution}
|
||||||
\vec \theta_\mathrm{Ridge} = (X^TX + \lambda I)^{-1}X^T \vec y.
|
\vec \theta_\mathrm{Ridge} = (X^TX + \lambda I)^{-1}X^T \vec y.
|
||||||
\end{equation}
|
\end{equation}
|
||||||
|
|
||||||
\subsubsection{Numerical Optimization}
|
\subsubsection{Numerical Optimization} \label{sec:numerical_optimization}
|
||||||
However, for Lasso regression there is no analytical solution and numerical methods must be used to obtain the optimal parameters $\vec \theta_\mathrm{Lasso}$. All numerical optimization methods are based on the idea of iteratively minimizing the cost function. Around the current guess for the parameters $\vec \theta_i$ the cost function can be approximated by a Taylor expansion as
|
However, for Lasso regression there is no analytical solution and numerical methods must be used to obtain the optimal parameters $\vec \theta_\mathrm{Lasso}$. All numerical optimization methods are based on the idea of iteratively minimizing the cost function. Around the current guess for the parameters $\vec \theta_i$ the cost function can be approximated by a Taylor expansion as \cite{hastie_elements_2009}
|
||||||
\begin{equation} \label{eq:taylor_expansion}
|
\begin{equation} \label{eq:taylor_expansion}
|
||||||
\begin{aligned}
|
\begin{aligned}
|
||||||
\cols(\vec \theta) &= \cols(\vec \theta_i) + \nabla \cols(\vec \theta_i)^T (\vec \theta - \vec \theta_i) \\&\quad+ \frac{1}{2} (\vec \theta - \vec \theta_i)^T H(\vec \theta_i) (\vec \theta - \vec \theta_i) \\&\quad+ \mathcal{O}\left(\norm{\vec \theta - \vec \theta_i}_2^3\right),
|
\cols(\vec \theta) &= \cols(\vec \theta_i) + \nabla \cols(\vec \theta_i)^T (\vec \theta - \vec \theta_i) \\&\quad+ \frac{1}{2} (\vec \theta - \vec \theta_i)^T H(\vec \theta_i) (\vec \theta - \vec \theta_i) \\&\quad+ \mathcal{O}\left(\norm{\vec \theta - \vec \theta_i}_2^3\right),
|
||||||
\end{aligned}
|
\end{aligned}
|
||||||
\end{equation}
|
\end{equation}
|
||||||
where $\nabla \cols(\vec \theta_i)$ is the gradient and $H(\vec \theta_i)$ the Hessian of the cost function at the current guess $\vec \theta_i$. Based on this approximation, the next guess for the parameters can be found by minimizing the quadratic approximation of the cost function. Since the calculation of the Hessian is computationally expensive, it is often neglected in practice. This leads to the gradient descent algorithm, where the next guess for the parameters is given by
|
where $\nabla \cols(\vec \theta_i)$ is the gradient and $H(\vec \theta_i)$ the Hessian of the cost function at the current guess $\vec \theta_i$. Based on this approximation, the next guess for the parameters can be found by minimizing the quadratic approximation of the cost function. Since the calculation of the Hessian is computationally expensive, it is often neglected in practice. This leads to the gradient descent algorithm, where the next guess for the parameters is given by \cite{goodfellow_deep_2016}
|
||||||
\begin{equation} \label{eq:gradient_descent}
|
\begin{equation} \label{eq:gradient_descent}
|
||||||
\vec \theta_{i+1} = \vec \theta_i - \eta \nabla \cols(\vec \theta_i),
|
\vec \theta_{i+1} = \vec \theta_i - \eta \nabla \cols(\vec \theta_i),
|
||||||
\end{equation}
|
\end{equation}
|
||||||
where $\eta$ is the learning rate that controls the step size of the update. The gradient descent algorithm can be further improved by using techniques such as momentum, Adagrad, RMSProp and Adam. These techniques adapt the learning rate based on the history of the gradients, leading to faster convergence and better performance.
|
where $\eta$ is the learning rate that controls the step size of the update. The gradient descent algorithm can be further improved by using techniques such as momentum, Adagrad, RMSProp and Adam. These techniques adapt the learning rate based on the history of the gradients, leading to faster convergence and better performance.
|
||||||
|
|
||||||
|
\paragraph{Momentum in Gradient Descent}
|
||||||
|
The momentum technique in gradient descent helps to accelerate the convergence of the optimization process by adding a fraction of the previous update to the current update. This helps to smooth out the updates and reduces oscillations in the optimization process. The update rule for gradient descent with momentum is given by
|
||||||
|
\begin{equation} \label{eq:momentum}
|
||||||
|
\begin{aligned}
|
||||||
|
v_{i} &= \delta \cdot \left( \theta_i - \theta_{i-1} - \eta \nabla \cols(\vec \theta_i) \right)\\
|
||||||
|
\vec \theta_{i+1} &= \vec \theta_i + v_i,
|
||||||
|
\end{aligned}
|
||||||
|
\end{equation}
|
||||||
|
where $v_i$ is the velocity vector that accumulates the gradients, $\beta$ is a hyperparameter that controls the decay rate of the velocity, and the other symbols have the same meaning as in the gradient descent update rule \cite{goodfellow_deep_2016,noauthor_stochastic_2025}.
|
||||||
|
|
||||||
\paragraph{Adagrad}
|
\paragraph{Adagrad}
|
||||||
The Adagrad algorithm adapts the learning rate for each parameter based on the history of the gradients. The update rule for Adagrad is given by
|
The Adagrad algorithm adapts the learning rate for each parameter based on the history of the gradients. The update rule for Adagrad is given by
|
||||||
\begin{equation} \label{eq:adagrad}
|
\begin{equation} \label{eq:adagrad}
|
||||||
@@ -83,7 +93,7 @@ The Adagrad algorithm adapts the learning rate for each parameter based on the h
|
|||||||
\vec \theta_{i+1} &= \vec \theta_i - \frac{\eta}{\sqrt{g_i + \epsilon}} \odot \nabla \cols(\vec \theta_i),
|
\vec \theta_{i+1} &= \vec \theta_i - \frac{\eta}{\sqrt{g_i + \epsilon}} \odot \nabla \cols(\vec \theta_i),
|
||||||
\end{aligned}
|
\end{aligned}
|
||||||
\end{equation}
|
\end{equation}
|
||||||
where $g_i$ is the accumulated squared gradient, $\odot$ denotes the element-wise product and $\epsilon$ is a small constant to prevent division by zero.
|
where $g_i$ is the accumulated squared gradient, $\odot$ denotes the element-wise product and $\epsilon$ is a small constant to prevent division by zero \cite{goodfellow_deep_2016,noauthor_stochastic_2025}.
|
||||||
|
|
||||||
|
|
||||||
\paragraph{RMSProp}
|
\paragraph{RMSProp}
|
||||||
@@ -94,7 +104,7 @@ The RMSProp algorithm is a modification of Adagrad that aims to reduce its aggre
|
|||||||
\vec \theta_{i+1} &= \vec \theta_i - \frac{\eta}{\sqrt{s_i + \epsilon}} \odot \nabla \cols(\vec \theta_i),
|
\vec \theta_{i+1} &= \vec \theta_i - \frac{\eta}{\sqrt{s_i + \epsilon}} \odot \nabla \cols(\vec \theta_i),
|
||||||
\end{aligned}
|
\end{aligned}
|
||||||
\end{equation}
|
\end{equation}
|
||||||
where $s_i$ is the moving average of the squared gradients, $\beta$ is a hyperparameter that controls the decay rate of the moving average, and the other symbols have the same meaning as in the Adagrad update rule.
|
where $s_i$ is the moving average of the squared gradients, $\beta$ is a hyperparameter that controls the decay rate of the moving average, and the other symbols have the same meaning as in the Adagrad update rule \cite{goodfellow_deep_2016,noauthor_stochastic_2025}.
|
||||||
|
|
||||||
\paragraph{Adam}
|
\paragraph{Adam}
|
||||||
The Adam (Adaptive Moment Estimation) algorithm combines the ideas of momentum and RMSProp. It maintains a moving average of both the gradients and the squared gradients, allowing it to adapt the learning rate for each parameter based on both the first and second moments of the gradients. The update rule for Adam is given by
|
The Adam (Adaptive Moment Estimation) algorithm combines the ideas of momentum and RMSProp. It maintains a moving average of both the gradients and the squared gradients, allowing it to adapt the learning rate for each parameter based on both the first and second moments of the gradients. The update rule for Adam is given by
|
||||||
@@ -105,23 +115,71 @@ The Adam (Adaptive Moment Estimation) algorithm combines the ideas of momentum a
|
|||||||
\vec \theta_{i+1} &= \vec \theta_i - \frac{\eta}{\sqrt{s_i + \epsilon}} \odot m_i,
|
\vec \theta_{i+1} &= \vec \theta_i - \frac{\eta}{\sqrt{s_i + \epsilon}} \odot m_i,
|
||||||
\end{aligned}
|
\end{aligned}
|
||||||
\end{equation}
|
\end{equation}
|
||||||
where $m_i$ is the moving average of the gradients, $s_i$ is the moving average of the squared gradients, $\beta_1$ and $\beta_2$ are hyperparameters that control the decay rates of the moving averages, and the other symbols have the same meaning as in the previous update rules.
|
where $m_i$ is the moving average of the gradients, $s_i$ is the moving average of the squared gradients, $\beta_1$ and $\beta_2$ are hyperparameters that control the decay rates of the moving averages, and the other symbols have the same meaning as in the previous update rules \cite{goodfellow_deep_2016,noauthor_stochastic_2025}.
|
||||||
|
|
||||||
\subsubsection{Stochastic Gradient Descent}
|
\subsubsection{Stochastic Gradient Descent}
|
||||||
As the computation of the gradient $\nabla \cols(\vec \theta_i)$ can be computationally expensive, especially for large datasets, a common approach is to use stochastic gradient descent (SGD). In SGD, the gradient is approximated by calculating it on a small subset of the data, called a mini-batch. This leads to a noisy estimate of the gradient, but it can significantly reduce the computational cost of each update step. The main challenge with SGD is to choose an appropriate mini-batch size. A small mini-batch size leads to a noisy estimate of the gradient, which can slow down convergence. On the other hand, a large mini-batch size can diminish the benefits of using SGD, as it approaches the full gradient descent computation cost. In practice the optimal mini-batch size depends on the specific problem and dataset.
|
As the computation of the gradient $\nabla \cols(\vec \theta_i)$ can be computationally expensive, especially for large datasets, a common approach is to use stochastic gradient descent (SGD). In SGD, the gradient is approximated by calculating it on a small subset of the data, called a mini-batch. This leads to a noisy estimate of the gradient, but it can significantly reduce the computational cost of each update step. The main challenge with SGD is to choose an appropriate mini-batch size. A small mini-batch size leads to a noisy estimate of the gradient, which can slow down convergence. On the other hand, a large mini-batch size can diminish the benefits of using SGD, as it approaches the full gradient descent computation cost. In practice the optimal mini-batch size depends on the specific problem and dataset\cite{goodfellow_deep_2016,elstner_lecture_2025}.
|
||||||
|
|
||||||
|
While the high variance of models with a high number of parameters can be counteracted with a large amount of data, the computational cost of fitting such models can become very high. This is especially relevant in the context of big data, where datasets can contain millions or even billions of samples. In such cases, it can be infeasible to use optimization algorithms that require the computation of the gradient over the entire dataset, such as ordinary gradient descent. For a dataset with $\N$ samples and a model with $p$ parameters, the computational cost of calculating the gradient
|
||||||
|
\begin{equation}
|
||||||
|
\nabla_\theta \cols = \frac{2}{N}\left( X^T X \vec \theta - X^T \vec{y} \right),
|
||||||
|
\end{equation}
|
||||||
|
is of order $\mathcal{O}(N p^2)$, as the matrix-vector products $X^T X$ needs to be computed. For very large datasets this can lead to very long computation times, as the entire dataset needs to be processed to compute the gradient. A common approach to counteract this problem is to use stochastic gradient descent (SGD) or mini-batch gradient descent, where the gradient is computed over a small subset of the data, called a batch. This leads to a significant reduction in computational cost, as the gradient can be computed in $\mathcal{O}(B p^2)$, where $B$ is the batch size. However, this comes at the cost of increased variance in the gradient estimate, as the gradient is only computed over a small subset of the data. This can lead to slower convergence and more oscillations in the optimization process.
|
||||||
|
|
||||||
|
While the decrease in computation time is significant, the result is even less black and white, when taking a closer look at the details.
|
||||||
|
While conventional GD can benefit largely by precomputing the matrix products $X^T X$ and $X^T \vec{y}$, this is not possible for SGD, as the batches change in every epoch. This leads to a significant increase in computation time per epoch for SGD compared to GD. However, as the number of epochs needed to reach convergence is usually much lower for SGD than for GD, the overall computation time can still be significantly lower for SGD. This is especially relevant for very large datasets, where the computational cost of GD can become prohibitive. Because while GD needs to process the entire dataset and thus the computational cost is of order $\mathcal{O}(N p^2)$, SGD only needs to process a small batch of size $B$ and thus the computational cost is of order $\mathcal{O}(B p^2)$ and invariant of the size of the dataset, given the batch size is large enough to estimate the complete gradient well enough. This leads to a significant reduction in computation time for very large datasets, as the computational cost of SGD does not increase with the size of the dataset\cite{goodfellow_deep_2016}.
|
||||||
|
|
||||||
\subsubsection{Resampling Methods}
|
\subsubsection{Resampling Methods}
|
||||||
To better understand the performance of a model, e.g. in terms of the bias-variance tradeoff, resampling methods such as bootstrapping and k-fold cross-validation can be used. These methods allow to estimate the performance of a model on unseen data by repeatedly splitting the data into training and test sets.
|
To better understand the performance of a model, e.g. in terms of the bias-variance tradeoff, resampling methods such as bootstrapping and k-fold cross-validation can be used. These methods allow to estimate the performance of a model on unseen data by repeatedly splitting the data into training and test sets.
|
||||||
\paragraph{Bootstrapping}
|
\paragraph{Bootstrapping}
|
||||||
In bootstrapping, multiple datasets are generated by randomly sampling the original training dataset with replacement. Each of these datasets is then used to train a model, and the performance of the model is evaluated on the unmodified test dataset. This process is repeated multiple times, and the performance metrics are averaged to obtain an estimate of the model's performance. Bootstrapping can be used to estimate the bias and variance of a model by analyzing the distribution of the performance metrics across the different bootstrap samples. The number of bootstrap samples is arbitrary, but a thousand samples are used in this work.
|
In bootstrapping, multiple datasets are generated by randomly sampling the original training dataset with replacement. Each of these datasets is then used to train a model, and the performance of the model is evaluated on the unmodified test dataset. This process is repeated multiple times, and the performance metrics are averaged to obtain an estimate of the model's performance. Bootstrapping can be used to estimate the bias and variance of a model by analyzing the distribution of the performance metrics across the different bootstrap samples\cite{hastie_elements_2009}. The number of bootstrap samples is arbitrary, but a thousand samples are used in this work.
|
||||||
\paragraph{k-Fold Cross-Validation}
|
\paragraph{k-Fold Cross-Validation}
|
||||||
In k-fold cross-validation, the original dataset is divided into $k$ equally sized folds. The model is then trained on $k-1$ folds and evaluated on the remaining fold. This process is repeated $k$ times, with each fold being used as the test set once. The performance metrics are then averaged across the $k$ iterations to obtain an estimate of the model's performance.
|
In k-fold cross-validation, the original dataset is divided into $k$ equally sized folds. The model is then trained on $k-1$ folds and evaluated on the remaining fold. This process is repeated $k$ times, with each fold being used as the test set once. The performance metrics are then averaged across the $k$ iterations to obtain an estimate of the model's performance\cite{elstner_lecture_2025}.
|
||||||
The process of k-fold cross-validation is illustrated in \cref{fig:crossvalidation}. For this work $k=5$ folds are used.
|
The process of k-fold cross-validation is illustrated in \cref{fig:crossvalidation}. For this work $k=5$ folds are used.
|
||||||
|
|
||||||
\begin{figure}
|
\begin{figure}
|
||||||
\centering
|
\centering
|
||||||
\includegraphics[width=\columnwidth]{../figures/crossvalidation.png}
|
\begin{tikzpicture}[scale=0.8]
|
||||||
\caption{Illustration of k-fold cross-validation with $k=5$. The dataset is divided into 5 folds, and the model is trained on 4 folds and evaluated on the remaining fold. This process is repeated 5 times, with each fold being used as the test set once.}
|
% parameters
|
||||||
|
\def\nfolds{5}
|
||||||
|
\def\squaresize{0.8}
|
||||||
|
\def\vspace{0.15} % extra space between rows (relative to squaresize)
|
||||||
|
|
||||||
|
% loop over folds
|
||||||
|
\foreach \row in {0,...,4} {
|
||||||
|
% vertical position with extra spacing
|
||||||
|
\pgfmathsetmacro{\y}{-(\row*(\squaresize+\vspace))}
|
||||||
|
\foreach \col in {0,...,4} {
|
||||||
|
% decide color
|
||||||
|
\ifnum\row=\col
|
||||||
|
\fill[turq] (\col*\squaresize, \y)
|
||||||
|
rectangle ++(\squaresize,\squaresize);
|
||||||
|
\else
|
||||||
|
\fill[white] (\col*\squaresize, \y)
|
||||||
|
rectangle ++(\squaresize,\squaresize);
|
||||||
|
\fi
|
||||||
|
% draw border
|
||||||
|
\draw[black] (\col*\squaresize, \y)
|
||||||
|
rectangle ++(\squaresize,\squaresize);
|
||||||
|
}
|
||||||
|
% label fold on the left
|
||||||
|
\node[left] at (-0.2, \y+0.5*\squaresize) {Fold \the\numexpr\row+1\relax};
|
||||||
|
% MSE_i label on the right
|
||||||
|
\node[right] at (5*\squaresize+0.2, \y+0.5*\squaresize) {$\mathrm{MSE}_{\the\numexpr\row+1\relax}$};
|
||||||
|
}
|
||||||
|
|
||||||
|
% curly brace summarizing MSEs
|
||||||
|
% compute top and bottom y positions for the brace
|
||||||
|
\pgfmathsetmacro{\ytop}{0.5*\squaresize}
|
||||||
|
\pgfmathsetmacro{\ybottom}{-(4*(\squaresize+\vspace))+0.5*\squaresize}
|
||||||
|
|
||||||
|
% curly brace summarizing MSEs
|
||||||
|
\draw[decorate,decoration={brace,amplitude=10pt}]
|
||||||
|
(6.8*\squaresize,\ytop) --
|
||||||
|
(6.8*\squaresize,\ybottom)
|
||||||
|
node[midway,xshift=1.5cm]{$\displaystyle \frac{1}{5}\sum_{i=1}^{5}\mathrm{MSE}_i$};
|
||||||
|
\end{tikzpicture}
|
||||||
|
\caption{Illustration of 5-fold cross-validation. Each row corresponds to one fold: the turquoise block marks the held-out test set while the white blocks form the training set. The models error $\mathrm{MSE}_i$ is computed for each fold, and the final performance is the average of all $\mathrm{MSE}_i$.}
|
||||||
\label{fig:crossvalidation}
|
\label{fig:crossvalidation}
|
||||||
\end{figure}
|
\end{figure}
|
||||||
|
|
||||||
@@ -131,14 +189,15 @@ The process of k-fold cross-validation is illustrated in \cref{fig:crossvalidati
|
|||||||
|
|
||||||
The methods described in the previous section have been implemented in Python. The implementation is structured in a modular way, allowing to easily switch between different optimization algorithms and resampling methods. The code is available on GitHub at \url{https://github.uio.no/larsbog/FYSSTK-Project1}. The following libraries have been used in the implementation:
|
The methods described in the previous section have been implemented in Python. The implementation is structured in a modular way, allowing to easily switch between different optimization algorithms and resampling methods. The code is available on GitHub at \url{https://github.uio.no/larsbog/FYSSTK-Project1}. The following libraries have been used in the implementation:
|
||||||
\begin{description}
|
\begin{description}
|
||||||
\item[\texttt{numpy}] It is a fundamental package for scientific computing in Python. It provides support for large, multi-dimensional arrays and matrices, along with a collection of mathematical functions to operate on these arrays efficiently. It is used in this project for all matrix and vector operations, as well as for generating random numbers and performing statistical calculations.
|
\item[\texttt{numpy}\cite{harris_array_2020}] It is a fundamental package for scientific computing in Python. It provides support for large, multi-dimensional arrays and matrices, along with a collection of mathematical functions to operate on these arrays efficiently. It is used in this project for all matrix and vector operations, as well as for generating random numbers and performing statistical calculations.
|
||||||
\item[\texttt{scikit-learn}] It is a machine learning library for Python that provides simple and efficient tools for data mining and data analysis. It includes implementations of various machine learning techniques, like data manipulation and performance evaluation metrics.
|
\item[\texttt{scikit-learn}\cite{pedregosa_scikit-learn_2011}] It is a machine learning library for Python that provides simple and efficient tools for data mining and data analysis. It includes implementations of various machine learning techniques, like data manipulation and performance evaluation metrics.
|
||||||
It is used in this project for rescaling the data, splitting the data into training and test sets, and for implementing the resampling methods. Furthermore, it provides the implementation of the mean squared error (MSE) metric used to evaluate the performance of the models.
|
It is used in this project for rescaling the data, splitting the data into training and test sets, and for implementing the resampling methods. Furthermore, it provides the implementation of the mean squared error (MSE) metric used to evaluate the performance of the models.
|
||||||
\item[\texttt{matplotlib}] It is a plotting library for Python that provides a wide range of tools for creating different types of plots and visualizations. It is used in this project to visualize the results of the different methods and to create plots for the report.
|
\item[\texttt{matplotlib}\cite{hunter_matplotlib_2007}] It is a plotting library for Python that provides a wide range of tools for creating different types of plots and visualizations. It is used in this project to visualize the results of the different methods and to create plots for the report.
|
||||||
\end{description}
|
\end{description}
|
||||||
|
|
||||||
\subsubsection{Other Software Tools}
|
\subsubsection{Other Software Tools}
|
||||||
The code has been developed using the \texttt{Visual Studio Code} editor, which provides a wide range of features for Python development, including syntax highlighting, code completion, and debugging tools. The version control system \texttt{Git} has been used to manage the codebase and track changes. The Large Language Model (LLM) \texttt{GitHub Copilot} has been used to assist in the development of the code by providing code suggestions and autocompletions based on the context of the code being written. All LLM-generated code has been reviewed and modified as necessary to ensure its correctness and suitability for the project.
|
The code has been developed using the \texttt{Visual Studio Code} editor, which provides a wide range of features for Python development, including syntax highlighting, code completion, and debugging tools. The version control system \texttt{Git} has been used to manage the codebase and track changes. The Large Language Model (LLM) \texttt{GitHub Copilot} has been used to assist in the development of the code by providing code suggestions and autocompletions based on the context of the code being written. All LLM-generated code has been reviewed and modified as necessary to ensure its correctness and suitability for the project.
|
||||||
|
The LLM ChatGPT-4 has been used to create \cref{fig:crossvalidation}. The LaTeX package \texttt{TikZ} has been used to create the figure. The LLM has been prompted to generate the TikZ code for the figure, which has then been modified as necessary to ensure its correctness and suitability for the report. The conversation can be found at the root of the GitHub repository.
|
||||||
|
|
||||||
|
|
||||||
\subsubsection{Implementation Details}
|
\subsubsection{Implementation Details}
|
||||||
@@ -148,5 +207,5 @@ The Runge function is defined as
|
|||||||
f(x) = \frac{1}{1 + 25x^2}, \quad x \in [-1, 1].
|
f(x) = \frac{1}{1 + 25x^2}, \quad x \in [-1, 1].
|
||||||
\end{equation}
|
\end{equation}
|
||||||
The $x$-values are sampled uniformly in the interval $[-1, 1]$, and the corresponding $y$-values are generated by adding Gaussian noise with mean 0 and standard deviation $\sigma = 1$, unless otherwise specified. To evaluate the performance in dependance of the degrees of freedom, the $x$-values are expanded into multiple polynomial features up to a specified degree: $X_i = \left\{x_i^0, x_i^1, \dots, x_i^p\right\}$. All features and outputs are then rescaled to have zero mean and unit variance using the \texttt{StandardScaler} from \texttt{scikit-learn}. The mean and variance of the training data are used for rescaling the test data to avoid data leakage.
|
The $x$-values are sampled uniformly in the interval $[-1, 1]$, and the corresponding $y$-values are generated by adding Gaussian noise with mean 0 and standard deviation $\sigma = 1$, unless otherwise specified. To evaluate the performance in dependance of the degrees of freedom, the $x$-values are expanded into multiple polynomial features up to a specified degree: $X_i = \left\{x_i^0, x_i^1, \dots, x_i^p\right\}$. All features and outputs are then rescaled to have zero mean and unit variance using the \texttt{StandardScaler} from \texttt{scikit-learn}. The mean and variance of the training data are used for rescaling the test data to avoid data leakage.
|
||||||
There are analytical implementations of the OLS and Ridge regression methods, as described in \cref{eq:ols_solution} and \cref{eq:ridge_solution}. Furthermore numerical optimizers are implemented from scratch based on a class inheritance structure to reduce code duplication. The base class \texttt{GradientDescent} implements a fitting procedure consisting of a precomputation step, an optimization procedure over multiple iterations and a memory of the cost function values. The different optimization algorithms, i.e. vanilla gradient descent, Adagrad, RMSProp and Adam, are implemented as subclasses that override the update rule for the parameters $\vec \theta$ in each iteration. The cost functions for OLS, Ridge and Lasso regression are also implemented as separate classes that provide methods to calculate the cost and its gradient. The stochastic gradient descent methods are implemented as further subclasses that modify the data used for fitting in each iteration. In the stochastic gradient descent method, the cost history property is furthermore modified to return the per epoch average of the cost function values. The resampling methods, i.e. bootstrapping and k-fold cross-validation, are implemented as methods with identical interfaces which provide multiple sets of training and test data. This allows to easily switch between the different resampling methods when evaluating the performance of a model.
|
There are analytical implementations of the OLS and Ridge regression methods, as described in \cref{eq:ols_solution} and \cref{eq:ridge_solution}. Furthermore numerical optimizers are implemented from scratch based on a class inheritance structure to reduce code duplication. The base class \texttt{GradientDescent} implements a fitting procedure consisting of a precomputation step, an optimization procedure over multiple iterations and a memory of the cost function values. The different optimization algorithms, i.e. vanilla gradient descent, momentum based gradient descent, Adagrad, RMSProp and Adam, are implemented as subclasses that override the update rule for the parameters $\vec \theta$ in each iteration. The cost functions for OLS, Ridge and Lasso regression are also implemented as separate classes that provide methods to calculate the cost and its gradient. The stochastic gradient descent methods are implemented as further subclasses that modify the data used for fitting in each iteration. In the stochastic gradient descent method, the cost history property is furthermore modified to return the per epoch average of the cost function values. The resampling methods, i.e. bootstrapping and k-fold cross-validation, are implemented as methods with identical interfaces which provide multiple sets of training and test data. This allows to easily switch between the different resampling methods when evaluating the performance of a model.
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1 @@
|
|||||||
|
With the increasing availability of large datasets and the growing complexity of models, there is an increasing need for efficient optimization techniques that can handle the computational challenges associated with big data. Especially in the context of deep learning, where models can have millions of parameters and require large amounts of data to train, the choice of optimization algorithm can have a significant impact on the performance of the model. As to limit the computational footprint of such datadriven methods, the search for more efficient optimization techniques is an ongoing field of research. There will also be further advances in the methods to choose hyperparameters of model training to make the usage of numerical models easier and more efficient to use.
|
||||||
@@ -47,12 +47,8 @@ With an increasing regularization strength $\lambda$ the parameter values are pu
|
|||||||
|
|
||||||
\subsection{Stochastic Gradient Descent}
|
\subsection{Stochastic Gradient Descent}
|
||||||
\label{sec:results_sgd}
|
\label{sec:results_sgd}
|
||||||
|
\Cref{fig:optimization_performance} shows the large decrease in computational cost, i.e. the decrease in computation time for \num{1000} epochs when using SGD instead of GD. The computation time decreases from \qty{14.82}{\s} in the case of OLS to \qty{3.71}{\s}. This is a decrease by over a factor of 3. For Ridge regression the computation time decreases from \qty{15.04}{\s} to \qty{4.29}{\s}. It is furthermore visible that the convergence of SGD is significantly faster during the first epochs as the parameters are updated more frequently. This leads to a significant decrease in the cost function during the first epochs, as the parameters are updated more frequently.
|
||||||
|
|
||||||
While the high variance of models with a high number of parameters can be counteracted with a large amount of data, the computational cost of fitting such models can become very high. This is especially relevant in the context of big data, where datasets can contain millions or even billions of samples. In such cases, it can be infeasible to use optimization algorithms that require the computation of the gradient over the entire dataset, such as ordinary gradient descent. For a dataset with $\N$ samples and a model with $p$ parameters, the computational cost of calculating the gradient
|
|
||||||
\begin{equation}
|
|
||||||
\nabla_\theta \cols = \frac{2}{N}\left( X^T X \vec \theta - X^T \vec{y} \right),
|
|
||||||
\end{equation}
|
|
||||||
is of order $\mathcal{O}(N p^2)$, as the matrix-vector products $X^T X$ needs to be computed. For very large datasets this can lead to very long computation times, as the entire dataset needs to be processed to compute the gradient. A common approach to counteract this problem is to use stochastic gradient descent (SGD) or mini-batch gradient descent, where the gradient is computed over a small subset of the data, called a batch. This leads to a significant reduction in computational cost, as the gradient can be computed in $\mathcal{O}(B p^2)$, where $B$ is the batch size. However, this comes at the cost of increased variance in the gradient estimate, as the gradient is only computed over a small subset of the data. This can lead to slower convergence and more oscillations in the optimization process.
|
|
||||||
|
|
||||||
\begin{figure}
|
\begin{figure}
|
||||||
\centering
|
\centering
|
||||||
@@ -61,16 +57,57 @@ is of order $\mathcal{O}(N p^2)$, as the matrix-vector products $X^T X$ needs to
|
|||||||
\label{fig:optimization_performance}
|
\label{fig:optimization_performance}
|
||||||
\end{figure}
|
\end{figure}
|
||||||
|
|
||||||
\Cref{fig:optimization_performance} shows the large decrease in computational cost, i.e. the decrease in computation time for \num{1000} epochs when using SGD instead of GD. While the decrease in computation time is significant, the result is even less black and white, when taking a closer look at the details.
|
|
||||||
While conventional GD can benefit largely by precomputing the matrix products $X^T X$ and $X^T \vec{y}$, this is not possible for SGD, as the batches change in every epoch. This leads to a significant increase in computation time per epoch for SGD compared to GD. However, as the number of epochs needed to reach convergence is usually much lower for SGD than for GD, the overall computation time can still be significantly lower for SGD. This is especially relevant for very large datasets, where the computational cost of GD can become prohibitive. Because while GD needs to process the entire dataset and thus the computational cost is of order $\mathcal{O}(N p^2)$, SGD only needs to process a small batch of size $B$ and thus the computational cost is of order $\mathcal{O}(B p^2)$ and invariant of the size of the dataset, given the batch size is large enough to estimate the complete gradient well enough. This leads to a significant reduction in computation time for very large datasets, as the computational cost of SGD does not increase with the size of the dataset.
|
|
||||||
|
|
||||||
A second aspect of SGD is visible in \cref{fig:optimization_performance}. While GD converges smoothly to the optimal solution, SGD shows significant oscillations around the optimal solution. This is a direct consequence of the increased variance in the gradient estimate, as the gradient is only computed over a small subset of the data. This can lead to slower convergence and even no convergence at all, if the batch size is too small or the noise in the data is too high. The intrinsic noise in the gradient estimate leads also to the issue, that SGD may be unable to reach an optimal solution with a large amount of noise in the training data. To counteract this problem, the noise in the dataset has been reduced by one order of magnitude compared to the previous sections. However, even with a reduced noise level, SGD is unable to reach the optimal solution. Further fine tuning of the new hyperparameters introduced by SGD, i.e. the batch size and the batches per epoch, could lead to better results. However, this is a non-trivial task and is highly dependent on the problem at hand. To arrive at a robust solution it might be necessary to increase the batch size until reaching the limit of full-batch GD, which would defeat the purpose of using SGD in the first place.
|
|
||||||
|
|
||||||
|
However a second aspect of SGD is visible in \cref{fig:optimization_performance}. While GD converges smoothly to the optimal solution, SGD shows significant oscillations around the optimal solution. This is a direct consequence of the increased variance in the gradient estimate, as the gradient is only computed over a small subset of the data. This can lead to slower convergence and even no convergence at all, if the batch size is too small or the noise in the data is too high. The intrinsic noise in the gradient estimate leads also to the issue, that SGD may be unable to reach an optimal solution with a large amount of noise in the training data. To counteract this problem, the noise in the dataset has been reduced by one order of magnitude compared to the previous sections. However, even with a reduced noise level, SGD is unable to reach the optimal solution. Further fine tuning of the new hyperparameters introduced by SGD, i.e. the batch size and the batches per epoch, could lead to better results. However, this is a non-trivial task and is highly dependent on the problem at hand. To arrive at a robust solution it might be necessary to increase the batch size until reaching the limit of full-batch GD, which would defeat the purpose of using SGD in the first place.
|
||||||
|
|
||||||
|
\subsection{Effect of Learning Rate on Convergence}
|
||||||
|
\label{sec:results_learning_rate}
|
||||||
|
A crucial hyperparameter in gradient-based optimization algorithms is the learning rate $\eta$. The learning rate determines the step size in the direction of the negative gradient and thus has a significant impact on the convergence behavior of the optimization algorithm. A learning rate that is too high can lead to divergence, as the optimization algorithm overshoots the optimal solution. On the other hand, a learning rate that is too low can lead to very slow convergence, as the optimization algorithm takes very small steps towards the optimal solution. In this section the effect of the learning rate on the convergence behavior of GD is studied.
|
||||||
|
|
||||||
|
\begin{figure}
|
||||||
|
\centering
|
||||||
|
\includegraphics[width=\columnwidth]{../figures/gradient_descent_convergence.pdf}
|
||||||
|
\caption{Convergence of GD for polynomial fits of degree 10 to noisy data from the Runge function using OLS and Ridge regression with $\lambda = \num{0.1}$. $\N=\num{80000}, N_\mathrm{iterations} = \num{1000}$ with different learning rates $\eta$.}
|
||||||
|
\label{fig:learning_rates}
|
||||||
|
\end{figure}
|
||||||
|
|
||||||
|
In \cref{fig:learning_rates} it is easily observable, that the convergence towards the minimum of the cost function accelerates with increasing learning rate. However, it is also visible that for very high learning rates, i.e. $\eta \geq 1$, the optimization algorithm diverges, as the cost function increases instead of decreasing. This is a direct consequence of the optimization algorithm overshooting the optimal solution and thus moving away from it instead of towards it. For very low learning rates, i.e. $\eta \leq 0.001$, the convergence is very slow, as the optimization algorithm takes very small steps towards the optimal solution. In such cases, it can be beneficial to increase the learning rate to speed up convergence. However, this comes at the risk of divergence, if the learning rate is increased too much. Thus, it is crucial to find a suitable learning rate that balances convergence speed and stability.
|
||||||
|
|
||||||
|
|
||||||
\subsection{Alternative Optimization Algorithms}
|
\subsection{Alternative Optimization Algorithms}
|
||||||
\label{sec:results_alt_opt_algos}
|
\label{sec:results_alt_opt_algos}
|
||||||
|
A different approach to limit the computational cost of fitting models with a high number of parameters is to use alternative optimization algorithms that reduce the number of gradient evaluations before reaching convergence. The optimized algorithms as introduced in \Cref{sec:numerical_optimization} are evaluated in regards to their convergence behavior on the same dataset and model.
|
||||||
|
|
||||||
|
\begin{figure}
|
||||||
|
\centering
|
||||||
|
\includegraphics[width=\columnwidth]{../figures/optimizer_comparison.pdf}
|
||||||
|
\caption{Minimization of OLS and Ridge regression cost functions using GD, Momentum GD, Adagrad, RMSProp and Adam optimization algorithms. $\N=\num{80000}, N_\mathrm{iterations}=250, p=10$ and learning rates $\eta_\mathrm{OLS} = \num{0.1}, \eta_\mathrm{Ridge} = \num{0.01}$.}
|
||||||
|
\label{fig:optimizer_comparison}
|
||||||
|
\end{figure}
|
||||||
|
|
||||||
|
\Cref{fig:optimizer_comparison} shows the convergence behavior of GD, Momentum GD, Adagrad, RMSProp and Adam optimization algorithms for OLS and Ridge regression. It is visible that all optimized algorithms converge faster than GD, with Adam showing the best performance. While GD is unable to reach perfect convergence within \num{250} iterations. Especially for the Ridge regression cost function, the optimized algorithms converge within a very short frame. On the contrary it is also visible that in multiple optimized algorithms there are stark oscillations in the convergence behavior. This is especially visible for the Momentum GD algorithm and the Adam algorithm. While these oscillations can be reduced by fine-tuning the hyperparameters of the algorithms, they are an intrinsic property of the algorithms and cannot be completely eliminated. However, as the overall convergence behavior is significantly improved compared to GD, these oscillations are usually not a problem in practice.
|
||||||
|
|
||||||
\subsection{Bias-Variance Tradeoff and Resampling Techniques}
|
\subsection{Bias-Variance Tradeoff and Resampling Techniques}
|
||||||
\label{sec:results_bias_variance}
|
\label{sec:results_bias_variance}
|
||||||
|
\begin{figure}
|
||||||
|
\centering
|
||||||
|
\includegraphics[width=\columnwidth]{../figures/bias_variance_tradeoff_combined_plot.pdf}
|
||||||
|
\caption{Bias-variance decomposition of the mean squared error (MSE) for polynomial fits of different degrees to noisy data from the Runge function using OLS. The Bias-Variance decomposition is calculated from bootstraped samples. The MSE is also calculated using 5-fold cross-validation and bootstrap resampling techniques. $\N=\num{80}, \Nt = \num{20}$.}
|
||||||
|
\label{fig:bias_variance_tradeoff}
|
||||||
|
\end{figure}
|
||||||
|
|
||||||
|
To evaluate the bias-variance tradeoff of the optimized models, the MSE is decomposed into its bias and variance components as introduced in \cref{sec:bias_variance_tradeoff}. The results are shown in \cref{fig:bias_variance_tradeoff}. To accurately depict the bias-variance tradeoff, there were two sets of test data used. To compute the bias, no noise was added to the test data, while for the rest of the computations the same noise level as in the training data was used. It is visible that for increasing polynomial degrees the variance increases, while the bias stays almost constant. This is a direct consequence of the increased flexibility of the model with increasing polynomial degree. To ensure numerical stability, Ridge regression with a very small regularization strength of $\lambda = \num{1e-10}$ has been used to compute the bias-variance decomposition in an approximation of OLS. This ensures that the matrix inversion in the computation of the model parameters is numerically stable, while the regularization term has almost no effect on the model performance.
|
||||||
|
The optimal model complexity is usually found at the point where the sum of bias and variance is minimal. Usually this coincides with the point of the minimum of the mean squared error. How the MSE behaves on average for $k$-fold cross-validation and bootstrap resampling techniques is also shown in \cref{fig:bias_variance_tradeoff}.
|
||||||
|
Both resampling techniques show a minimum in the MSE for for the polynomial degree of 2. The crossvalidation technique shows a much more constant MSE value over the entire range of polynomial degrees, while the bootstrap technique shows a more pronounced minimum. This indicates that the bootstrap technique is more sensitive to the choice of polynomial degree and thus provides a better estimate in the case of hyperparameter tuning at the cost of a much higher computational cost.
|
||||||
|
|
||||||
|
This trend of optimal model complexity being achieved somewhere between the high bias case of low parameter count and the high variance case of high parameter count is a general property of machine learning models and is not limited to polynomial regression with OLS cost functions. As seen in \cref{fig:kfold_mse_comp} this trend of a minimum in the MSE is also observable for Ridge and Lasso regression.
|
||||||
|
|
||||||
|
\begin{figure}
|
||||||
|
\centering
|
||||||
|
\includegraphics[width=\columnwidth]{../figures/kfold_mse_comparison_per_cost_function.pdf}
|
||||||
|
\caption{Mean squared error (MSE) for polynomial fits of different degrees to noisy data from the Runge function using OLS, Ridge and Lasso regression cost functions with 5-fold cross-validation resampling technique. $\N=\num{240}, \Nt = \num{60}$ and $\lambda = \num{0.1}$ for Ridge and Lasso regression.}
|
||||||
|
\label{fig:kfold_mse_comp}
|
||||||
|
\end{figure}
|
||||||
|
|
||||||
|
It is also observable, that the minimum of the MSE is shifted towards higher polynomial degrees for Ridge and Lasso regression with a slower increase in the mean squared error above the optimal polynomial degree compared to OLS regression. This is a direct consequence of the regularization term in the cost function, which decreases the effective degrees of freedom of the model. Thus, a model with a higher number of parameters is needed to achieve the same flexibility as a model with a lower number of parameters without regularization. This effect is especially strong for Lasso regression, where the regularization term can lead to some parameters being set to zero, effectively removing them from the model. This leads to a significant reduction in the effective degrees of freedom of the model and thus a shift of the optimal polynomial degree towards higher values. The reduction in effective degrees of freedom always depends on the data at hand and the chosen regularization strength $\lambda$ and thus cannot be quantified in a general way.
|
||||||
|
|||||||
@@ -0,0 +1,442 @@
|
|||||||
|
{
|
||||||
|
"title": "Tikz figure for cross-validation",
|
||||||
|
"create_time": 1758533017.029746,
|
||||||
|
"update_time": 1758694769.6012,
|
||||||
|
"mapping": {
|
||||||
|
"client-created-root": {
|
||||||
|
"id": "client-created-root",
|
||||||
|
"message": null,
|
||||||
|
"parent": null,
|
||||||
|
"children": ["62419da9-6897-4c64-bc3d-e4e0210c290a"]
|
||||||
|
},
|
||||||
|
"62419da9-6897-4c64-bc3d-e4e0210c290a": {
|
||||||
|
"id": "62419da9-6897-4c64-bc3d-e4e0210c290a",
|
||||||
|
"message": {
|
||||||
|
"id": "62419da9-6897-4c64-bc3d-e4e0210c290a",
|
||||||
|
"author": { "role": "system", "name": null, "metadata": {} },
|
||||||
|
"create_time": null,
|
||||||
|
"update_time": null,
|
||||||
|
"content": { "content_type": "text", "parts": [""] },
|
||||||
|
"status": "finished_successfully",
|
||||||
|
"end_turn": true,
|
||||||
|
"weight": 0.0,
|
||||||
|
"metadata": { "is_visually_hidden_from_conversation": true },
|
||||||
|
"recipient": "all",
|
||||||
|
"channel": null
|
||||||
|
},
|
||||||
|
"parent": "client-created-root",
|
||||||
|
"children": ["fd0f2eef-ff5d-4de7-af0f-6a64a680b7de"]
|
||||||
|
},
|
||||||
|
"fd0f2eef-ff5d-4de7-af0f-6a64a680b7de": {
|
||||||
|
"id": "fd0f2eef-ff5d-4de7-af0f-6a64a680b7de",
|
||||||
|
"message": {
|
||||||
|
"id": "fd0f2eef-ff5d-4de7-af0f-6a64a680b7de",
|
||||||
|
"author": { "role": "system", "name": null, "metadata": {} },
|
||||||
|
"create_time": 1758532934.1949592,
|
||||||
|
"update_time": null,
|
||||||
|
"content": { "content_type": "text", "parts": [""] },
|
||||||
|
"status": "finished_successfully",
|
||||||
|
"end_turn": null,
|
||||||
|
"weight": 1.0,
|
||||||
|
"metadata": { "is_visually_hidden_from_conversation": true },
|
||||||
|
"recipient": "all",
|
||||||
|
"channel": null
|
||||||
|
},
|
||||||
|
"parent": "62419da9-6897-4c64-bc3d-e4e0210c290a",
|
||||||
|
"children": ["3becc95e-385d-47dd-ba12-b551b85c93fe"]
|
||||||
|
},
|
||||||
|
"3becc95e-385d-47dd-ba12-b551b85c93fe": {
|
||||||
|
"id": "3becc95e-385d-47dd-ba12-b551b85c93fe",
|
||||||
|
"message": {
|
||||||
|
"id": "3becc95e-385d-47dd-ba12-b551b85c93fe",
|
||||||
|
"author": { "role": "user", "name": null, "metadata": {} },
|
||||||
|
"create_time": null,
|
||||||
|
"update_time": null,
|
||||||
|
"content": {
|
||||||
|
"content_type": "user_editable_context",
|
||||||
|
"user_profile": "The user provided the following information about themselves. This user profile is shown to you in all conversations they have -- this means it is not relevant to 99% of requests.\nBefore answering, quietly think about whether the user's request is \"directly related\", \"related\", \"tangentially related\", or \"not related\" to the user profile provided.\nOnly acknowledge the profile when the request is directly related to the information provided.\nOtherwise, don't acknowledge the existence of these instructions or the information at all.\nUser profile:\n```Preferred name: Lars\nRole: Student\nOther Information: Interested in Photography, Networking, Python, Data Science, Particle Physics, Electronics. Critical Thinker.\n```",
|
||||||
|
"user_instructions": "The user provided the additional info about how they would like you to respond:\n```Use quick and clever humor when appropriate. Get right to the point. Be innovative and think outside the box.```"
|
||||||
|
},
|
||||||
|
"status": "finished_successfully",
|
||||||
|
"end_turn": null,
|
||||||
|
"weight": 1.0,
|
||||||
|
"metadata": {
|
||||||
|
"is_visually_hidden_from_conversation": true,
|
||||||
|
"user_context_message_data": {
|
||||||
|
"about_user_message": "Preferred name: Lars\nRole: Student\nOther Information: Interested in Photography, Networking, Python, Data Science, Particle Physics, Electronics. Critical Thinker.\n",
|
||||||
|
"about_model_message": "Use quick and clever humor when appropriate. Get right to the point. Be innovative and think outside the box."
|
||||||
|
},
|
||||||
|
"is_user_system_message": true
|
||||||
|
},
|
||||||
|
"recipient": "all",
|
||||||
|
"channel": null
|
||||||
|
},
|
||||||
|
"parent": "fd0f2eef-ff5d-4de7-af0f-6a64a680b7de",
|
||||||
|
"children": ["b2f9f24f-13ae-4d3d-b63f-2aa59aebc36d"]
|
||||||
|
},
|
||||||
|
"b2f9f24f-13ae-4d3d-b63f-2aa59aebc36d": {
|
||||||
|
"id": "b2f9f24f-13ae-4d3d-b63f-2aa59aebc36d",
|
||||||
|
"message": {
|
||||||
|
"id": "b2f9f24f-13ae-4d3d-b63f-2aa59aebc36d",
|
||||||
|
"author": { "role": "user", "name": null, "metadata": {} },
|
||||||
|
"create_time": 1758533016.597,
|
||||||
|
"update_time": null,
|
||||||
|
"content": {
|
||||||
|
"content_type": "text",
|
||||||
|
"parts": [
|
||||||
|
"Create a tikz figure to explain 5-fold-crossvalidation.\n\n5 rows of 5 squares each. 4 white one in turqouise, where the colored square moves through the places for each row. Every row represents one fold."
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"status": "finished_successfully",
|
||||||
|
"end_turn": null,
|
||||||
|
"weight": 1.0,
|
||||||
|
"metadata": {
|
||||||
|
"selected_github_repos": [],
|
||||||
|
"serialization_metadata": { "custom_symbol_offsets": [] },
|
||||||
|
"request_id": "9830be9aed763181-MXP",
|
||||||
|
"message_source": null,
|
||||||
|
"timestamp_": "absolute"
|
||||||
|
},
|
||||||
|
"recipient": "all",
|
||||||
|
"channel": null
|
||||||
|
},
|
||||||
|
"parent": "3becc95e-385d-47dd-ba12-b551b85c93fe",
|
||||||
|
"children": ["6a10ac9c-c8a9-45de-9d1a-acf16d2e7f12"]
|
||||||
|
},
|
||||||
|
"6a10ac9c-c8a9-45de-9d1a-acf16d2e7f12": {
|
||||||
|
"id": "6a10ac9c-c8a9-45de-9d1a-acf16d2e7f12",
|
||||||
|
"message": {
|
||||||
|
"id": "6a10ac9c-c8a9-45de-9d1a-acf16d2e7f12",
|
||||||
|
"author": { "role": "assistant", "name": null, "metadata": {} },
|
||||||
|
"create_time": 1758533016.9494157,
|
||||||
|
"update_time": null,
|
||||||
|
"content": { "content_type": "text", "parts": [""] },
|
||||||
|
"status": "finished_successfully",
|
||||||
|
"end_turn": null,
|
||||||
|
"weight": 1.0,
|
||||||
|
"metadata": {
|
||||||
|
"request_id": "9830be9aed763181-MXP",
|
||||||
|
"message_type": "next",
|
||||||
|
"model_slug": "gpt-5",
|
||||||
|
"default_model_slug": "auto",
|
||||||
|
"parent_id": "b2f9f24f-13ae-4d3d-b63f-2aa59aebc36d",
|
||||||
|
"timestamp_": "absolute"
|
||||||
|
},
|
||||||
|
"recipient": "all",
|
||||||
|
"channel": null
|
||||||
|
},
|
||||||
|
"parent": "b2f9f24f-13ae-4d3d-b63f-2aa59aebc36d",
|
||||||
|
"children": ["d2ada739-5007-46fb-886d-df96115657eb"]
|
||||||
|
},
|
||||||
|
"d2ada739-5007-46fb-886d-df96115657eb": {
|
||||||
|
"id": "d2ada739-5007-46fb-886d-df96115657eb",
|
||||||
|
"message": {
|
||||||
|
"id": "d2ada739-5007-46fb-886d-df96115657eb",
|
||||||
|
"author": { "role": "assistant", "name": null, "metadata": {} },
|
||||||
|
"create_time": 1758533016.968473,
|
||||||
|
"update_time": null,
|
||||||
|
"content": {
|
||||||
|
"content_type": "text",
|
||||||
|
"parts": [
|
||||||
|
"Here\u2019s a minimal LaTeX example with TikZ that draws exactly what you described: a 5\u00d75 grid where each row corresponds to a fold, and the turquoise square marks the held-out fold for that row. \n\n```latex\n\\documentclass[tikz]{standalone}\n\\usepackage{xcolor}\n\\definecolor{turq}{RGB}{64,224,208}\n\n\\begin{document}\n\\begin{tikzpicture}[scale=0.8]\n % parameters\n \\def\\nfolds{5}\n \\def\\squaresize{0.8}\n\n % loop over folds\n \\foreach \\row in {0,...,4} {\n \\foreach \\col in {0,...,4} {\n % decide color: turquoise if col=row, else white\n \\ifnum\\row=\\col\n \\fill[turq] (\\col*\\squaresize, -\\row*\\squaresize) \n rectangle ++(\\squaresize,\\squaresize);\n \\else\n \\fill[white] (\\col*\\squaresize, -\\row*\\squaresize) \n rectangle ++(\\squaresize,\\squaresize);\n \\fi\n % draw border\n \\draw[black] (\\col*\\squaresize, -\\row*\\squaresize) \n rectangle ++(\\squaresize,\\squaresize);\n }\n % label row as Fold number\n \\node[left] at (-0.2,-\\row*\\squaresize+0.4*\\squaresize) {Fold \\the\\numexpr\\row+1\\relax};\n }\n\\end{tikzpicture}\n\\end{document}\n```\n\nThis produces:\n\n- 5 rows (Fold 1 to Fold 5). \n- 5 squares per row. \n- Exactly one turquoise (`turq`) square per row, shifting its position across columns. \n- Labels on the left for clarity. \n\nYou can easily tweak `\\squaresize`, the turquoise color definition, or remove the labels if you want a cleaner diagram. \n\nDo you want me to also add arrows or braces to visually emphasize \u201ctraining set\u201d vs \u201cvalidation set\u201d in each row?"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"status": "finished_successfully",
|
||||||
|
"end_turn": true,
|
||||||
|
"weight": 1.0,
|
||||||
|
"metadata": {
|
||||||
|
"finish_details": { "type": "stop", "stop_tokens": [200002] },
|
||||||
|
"is_complete": true,
|
||||||
|
"citations": [],
|
||||||
|
"content_references": [],
|
||||||
|
"request_id": "9830be9aed763181-MXP",
|
||||||
|
"message_type": "next",
|
||||||
|
"model_slug": "gpt-5",
|
||||||
|
"default_model_slug": "auto",
|
||||||
|
"parent_id": "6a10ac9c-c8a9-45de-9d1a-acf16d2e7f12",
|
||||||
|
"timestamp_": "absolute"
|
||||||
|
},
|
||||||
|
"recipient": "all",
|
||||||
|
"channel": null
|
||||||
|
},
|
||||||
|
"parent": "6a10ac9c-c8a9-45de-9d1a-acf16d2e7f12",
|
||||||
|
"children": ["3610d3d8-8dfb-4b3b-ac57-5e045efedc3c"]
|
||||||
|
},
|
||||||
|
"3610d3d8-8dfb-4b3b-ac57-5e045efedc3c": {
|
||||||
|
"id": "3610d3d8-8dfb-4b3b-ac57-5e045efedc3c",
|
||||||
|
"message": {
|
||||||
|
"id": "3610d3d8-8dfb-4b3b-ac57-5e045efedc3c",
|
||||||
|
"author": { "role": "user", "name": null, "metadata": {} },
|
||||||
|
"create_time": 1758533188.733,
|
||||||
|
"update_time": null,
|
||||||
|
"content": {
|
||||||
|
"content_type": "text",
|
||||||
|
"parts": [
|
||||||
|
"Add a vertical space of 0.15 squarespace between the rows. In the end of each row write MSE_i where i is the fold number and then a vertical curly brace summarizing all mses into the average of the mses"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"status": "finished_successfully",
|
||||||
|
"end_turn": null,
|
||||||
|
"weight": 1.0,
|
||||||
|
"metadata": {
|
||||||
|
"selected_github_repos": [],
|
||||||
|
"serialization_metadata": { "custom_symbol_offsets": [] },
|
||||||
|
"request_id": "9830c2ceecfa3181-MXP",
|
||||||
|
"message_source": null,
|
||||||
|
"timestamp_": "absolute"
|
||||||
|
},
|
||||||
|
"recipient": "all",
|
||||||
|
"channel": null
|
||||||
|
},
|
||||||
|
"parent": "d2ada739-5007-46fb-886d-df96115657eb",
|
||||||
|
"children": ["16b2e514-fdfe-4e92-b6cd-c4a0c217a415"]
|
||||||
|
},
|
||||||
|
"16b2e514-fdfe-4e92-b6cd-c4a0c217a415": {
|
||||||
|
"id": "16b2e514-fdfe-4e92-b6cd-c4a0c217a415",
|
||||||
|
"message": {
|
||||||
|
"id": "16b2e514-fdfe-4e92-b6cd-c4a0c217a415",
|
||||||
|
"author": { "role": "assistant", "name": null, "metadata": {} },
|
||||||
|
"create_time": 1758533189.143814,
|
||||||
|
"update_time": null,
|
||||||
|
"content": {
|
||||||
|
"content_type": "text",
|
||||||
|
"parts": [
|
||||||
|
"Perfect \u2014 let\u2019s pimp this cross-validation diagram a bit. I\u2019ll: \n\n* Add 0.15 \u00d7 `squaresize` vertical spacing between rows. \n* Put `MSE_i` at the right end of each row. \n* Collect them with a big brace on the right, pointing to the average. \n\nHere\u2019s the full TikZ code: \n\n```latex\n\\documentclass[tikz]{standalone}\n\\usepackage{xcolor}\n\\definecolor{turq}{RGB}{64,224,208}\n\n\\begin{document}\n\\begin{tikzpicture}[scale=0.8]\n % parameters\n \\def\\nfolds{5}\n \\def\\squaresize{0.8}\n \\def\\vspace{0.15} % extra space between rows (relative to squaresize)\n\n % loop over folds\n \\foreach \\row in {0,...,4} {\n % vertical position with extra spacing\n \\pgfmathsetmacro{\\y}{-(\\row*(\\squaresize+\\vspace))}\n \\foreach \\col in {0,...,4} {\n % decide color\n \\ifnum\\row=\\col\n \\fill[turq] (\\col*\\squaresize, \\y) \n rectangle ++(\\squaresize,\\squaresize);\n \\else\n \\fill[white] (\\col*\\squaresize, \\y) \n rectangle ++(\\squaresize,\\squaresize);\n \\fi\n % draw border\n \\draw[black] (\\col*\\squaresize, \\y) \n rectangle ++(\\squaresize,\\squaresize);\n }\n % label fold on the left\n \\node[left] at (-0.2, \\y+0.5*\\squaresize) {Fold \\the\\numexpr\\row+1\\relax};\n % MSE_i label on the right\n \\node[right] at (5*\\squaresize+0.2, \\y+0.5*\\squaresize) {$\\mathrm{MSE}_{\\the\\numexpr\\row+1\\relax}$};\n }\n\n % curly brace summarizing MSEs\n \\draw[decorate,decoration={brace,amplitude=10pt}] \n (5.8*\\squaresize,0.5*\\squaresize) -- \n (5.8*\\squaresize,-4*(\\squaresize+\\vspace)+0.5*\\squaresize)\n node[midway,xshift=2.0cm]{$\\displaystyle \\frac{1}{5}\\sum_{i=1}^{5}\\mathrm{MSE}_i$};\n\\end{tikzpicture}\n\\end{document}\n```\n\nThis will render: \n\n- Five staggered rows with the turquoise square marching across. \n- `MSE\u2081, \u2026, MSE\u2085` at the row ends. \n- A vertical brace on the right summarizing them into their average. \n\nWant me to also annotate explicitly which part is \u201ctraining\u201d vs \u201cvalidation\u201d inside each row, so it doubles as a teaching slide?"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"status": "finished_successfully",
|
||||||
|
"end_turn": true,
|
||||||
|
"weight": 1.0,
|
||||||
|
"metadata": {
|
||||||
|
"finish_details": { "type": "stop", "stop_tokens": [200002] },
|
||||||
|
"is_complete": true,
|
||||||
|
"citations": [],
|
||||||
|
"content_references": [],
|
||||||
|
"request_id": "9830c2ceecfa3181-MXP",
|
||||||
|
"message_type": "next",
|
||||||
|
"model_slug": "gpt-5",
|
||||||
|
"default_model_slug": "auto",
|
||||||
|
"parent_id": "3610d3d8-8dfb-4b3b-ac57-5e045efedc3c",
|
||||||
|
"timestamp_": "absolute"
|
||||||
|
},
|
||||||
|
"recipient": "all",
|
||||||
|
"channel": null
|
||||||
|
},
|
||||||
|
"parent": "3610d3d8-8dfb-4b3b-ac57-5e045efedc3c",
|
||||||
|
"children": ["f05d8662-5e8a-4e5c-bf64-7847c8a2cb89"]
|
||||||
|
},
|
||||||
|
"f05d8662-5e8a-4e5c-bf64-7847c8a2cb89": {
|
||||||
|
"id": "f05d8662-5e8a-4e5c-bf64-7847c8a2cb89",
|
||||||
|
"message": {
|
||||||
|
"id": "f05d8662-5e8a-4e5c-bf64-7847c8a2cb89",
|
||||||
|
"author": { "role": "user", "name": null, "metadata": {} },
|
||||||
|
"create_time": 1758533441.129,
|
||||||
|
"update_time": null,
|
||||||
|
"content": {
|
||||||
|
"content_type": "text",
|
||||||
|
"parts": [
|
||||||
|
"Package pgfkeys: I do not know the key '/pgf/decoration/.expanded' and I am going to ignore it. Perhaps you misspelled it.LaTeX"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"status": "finished_successfully",
|
||||||
|
"end_turn": null,
|
||||||
|
"weight": 1.0,
|
||||||
|
"metadata": {
|
||||||
|
"selected_github_repos": [],
|
||||||
|
"serialization_metadata": { "custom_symbol_offsets": [] },
|
||||||
|
"request_id": "9830c8f8588db518-OSL",
|
||||||
|
"message_source": null,
|
||||||
|
"timestamp_": "absolute"
|
||||||
|
},
|
||||||
|
"recipient": "all",
|
||||||
|
"channel": null
|
||||||
|
},
|
||||||
|
"parent": "16b2e514-fdfe-4e92-b6cd-c4a0c217a415",
|
||||||
|
"children": ["45d98e06-ed62-4b57-84c6-4fe2c433071a"]
|
||||||
|
},
|
||||||
|
"45d98e06-ed62-4b57-84c6-4fe2c433071a": {
|
||||||
|
"id": "45d98e06-ed62-4b57-84c6-4fe2c433071a",
|
||||||
|
"message": {
|
||||||
|
"id": "45d98e06-ed62-4b57-84c6-4fe2c433071a",
|
||||||
|
"author": { "role": "assistant", "name": null, "metadata": {} },
|
||||||
|
"create_time": 1758533442.068548,
|
||||||
|
"update_time": null,
|
||||||
|
"content": {
|
||||||
|
"content_type": "text",
|
||||||
|
"parts": [
|
||||||
|
"That error happens when TikZ doesn\u2019t know about *decorations*. By default, braces are part of the `decorations.pathreplacing` library, which you need to load explicitly. Without it, TikZ grumbles with that message. \n\nHere\u2019s a fixed version that will compile cleanly: \n\n```latex\n\\documentclass[tikz]{standalone}\n\\usepackage{xcolor}\n\\usetikzlibrary{decorations.pathreplacing} % <-- needed for the brace\n\n\\definecolor{turq}{RGB}{64,224,208}\n\n\\begin{document}\n\\begin{tikzpicture}[scale=0.8]\n % parameters\n \\def\\nfolds{5}\n \\def\\squaresize{0.8}\n \\def\\vspace{0.15} % extra space between rows (relative to squaresize)\n\n % loop over folds\n \\foreach \\row in {0,...,4} {\n % vertical position with extra spacing\n \\pgfmathsetmacro{\\y}{-(\\row*(\\squaresize+\\vspace))}\n \\foreach \\col in {0,...,4} {\n % decide color\n \\ifnum\\row=\\col\n \\fill[turq] (\\col*\\squaresize, \\y) \n rectangle ++(\\squaresize,\\squaresize);\n \\else\n \\fill[white] (\\col*\\squaresize, \\y) \n rectangle ++(\\squaresize,\\squaresize);\n \\fi\n % draw border\n \\draw[black] (\\col*\\squaresize, \\y) \n rectangle ++(\\squaresize,\\squaresize);\n }\n % label fold on the left\n \\node[left] at (-0.2, \\y+0.5*\\squaresize) {Fold \\the\\numexpr\\row+1\\relax};\n % MSE_i label on the right\n \\node[right] at (5*\\squaresize+0.2, \\y+0.5*\\squaresize) {$\\mathrm{MSE}_{\\the\\numexpr\\row+1\\relax}$};\n }\n\n % curly brace summarizing MSEs\n \\draw[decorate,decoration={brace,amplitude=10pt}] \n (5.8*\\squaresize,0.5*\\squaresize) -- \n (5.8*\\squaresize,-4*(\\squaresize+\\vspace)+0.5*\\squaresize)\n node[midway,xshift=2.0cm]{$\\displaystyle \\frac{1}{5}\\sum_{i=1}^{5}\\mathrm{MSE}_i$};\n\\end{tikzpicture}\n\\end{document}\n```\n\nNow TikZ knows how to draw that brace. \n\nWould you like me to also make the brace point *into* the `MSE_i` labels (so it visually groups them), or should it just sit neatly outside?"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"status": "finished_successfully",
|
||||||
|
"end_turn": true,
|
||||||
|
"weight": 1.0,
|
||||||
|
"metadata": {
|
||||||
|
"finish_details": { "type": "stop", "stop_tokens": [200002] },
|
||||||
|
"is_complete": true,
|
||||||
|
"citations": [],
|
||||||
|
"content_references": [],
|
||||||
|
"request_id": "9830c8f8588db518-OSL",
|
||||||
|
"message_type": "next",
|
||||||
|
"model_slug": "gpt-5",
|
||||||
|
"default_model_slug": "auto",
|
||||||
|
"parent_id": "f05d8662-5e8a-4e5c-bf64-7847c8a2cb89",
|
||||||
|
"timestamp_": "absolute"
|
||||||
|
},
|
||||||
|
"recipient": "all",
|
||||||
|
"channel": null
|
||||||
|
},
|
||||||
|
"parent": "f05d8662-5e8a-4e5c-bf64-7847c8a2cb89",
|
||||||
|
"children": ["e1f63acf-2ae2-4bca-9b2b-508b705c07c0"]
|
||||||
|
},
|
||||||
|
"e1f63acf-2ae2-4bca-9b2b-508b705c07c0": {
|
||||||
|
"id": "e1f63acf-2ae2-4bca-9b2b-508b705c07c0",
|
||||||
|
"message": {
|
||||||
|
"id": "e1f63acf-2ae2-4bca-9b2b-508b705c07c0",
|
||||||
|
"author": { "role": "user", "name": null, "metadata": {} },
|
||||||
|
"create_time": 1758533512.121,
|
||||||
|
"update_time": null,
|
||||||
|
"content": {
|
||||||
|
"content_type": "text",
|
||||||
|
"parts": [
|
||||||
|
" (5.8*\\squaresize,-4*(\\squaresize+\\vspace)+0.5*\\squaresize)\n\nCannot parse this coordinate"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"status": "finished_successfully",
|
||||||
|
"end_turn": null,
|
||||||
|
"weight": 1.0,
|
||||||
|
"metadata": {
|
||||||
|
"selected_github_repos": [],
|
||||||
|
"serialization_metadata": { "custom_symbol_offsets": [] },
|
||||||
|
"request_id": "9830cab47b4db518-MXP",
|
||||||
|
"message_source": null,
|
||||||
|
"timestamp_": "absolute"
|
||||||
|
},
|
||||||
|
"recipient": "all",
|
||||||
|
"channel": null
|
||||||
|
},
|
||||||
|
"parent": "45d98e06-ed62-4b57-84c6-4fe2c433071a",
|
||||||
|
"children": ["f88fa5f2-0361-4186-a8ff-cfcc7f3328ac"]
|
||||||
|
},
|
||||||
|
"f88fa5f2-0361-4186-a8ff-cfcc7f3328ac": {
|
||||||
|
"id": "f88fa5f2-0361-4186-a8ff-cfcc7f3328ac",
|
||||||
|
"message": {
|
||||||
|
"id": "f88fa5f2-0361-4186-a8ff-cfcc7f3328ac",
|
||||||
|
"author": { "role": "assistant", "name": null, "metadata": {} },
|
||||||
|
"create_time": 1758533512.668357,
|
||||||
|
"update_time": null,
|
||||||
|
"content": {
|
||||||
|
"content_type": "text",
|
||||||
|
"parts": [
|
||||||
|
"Ah yes \u2014 TikZ doesn\u2019t parse arithmetic like `-4*(\\squaresize+\\vspace)` directly inside coordinates. You have to precompute with `\\pgfmathsetmacro` or wrap the arithmetic in `{}` with `\\dimexpr`. \n\nHere\u2019s a corrected version that will compile cleanly: \n\n```latex\n\\documentclass[tikz]{standalone}\n\\usepackage{xcolor}\n\\usetikzlibrary{decorations.pathreplacing} % for braces\n\n\\definecolor{turq}{RGB}{64,224,208}\n\n\\begin{document}\n\\begin{tikzpicture}[scale=0.8]\n % parameters\n \\def\\nfolds{5}\n \\def\\squaresize{0.8}\n \\def\\vspace{0.15} % extra space between rows (relative to squaresize)\n\n % loop over folds\n \\foreach \\row in {0,...,4} {\n % vertical position with extra spacing\n \\pgfmathsetmacro{\\y}{-(\\row*(\\squaresize+\\vspace))}\n \\foreach \\col in {0,...,4} {\n % decide color\n \\ifnum\\row=\\col\n \\fill[turq] (\\col*\\squaresize, \\y) \n rectangle ++(\\squaresize,\\squaresize);\n \\else\n \\fill[white] (\\col*\\squaresize, \\y) \n rectangle ++(\\squaresize,\\squaresize);\n \\fi\n % draw border\n \\draw[black] (\\col*\\squaresize, \\y) \n rectangle ++(\\squaresize,\\squaresize);\n }\n % label fold on the left\n \\node[left] at (-0.2, \\y+0.5*\\squaresize) {Fold \\the\\numexpr\\row+1\\relax};\n % MSE_i label on the right\n \\node[right] at (5*\\squaresize+0.2, \\y+0.5*\\squaresize) {$\\mathrm{MSE}_{\\the\\numexpr\\row+1\\relax}$};\n }\n\n % compute top and bottom y positions for the brace\n \\pgfmathsetmacro{\\ytop}{0.5*\\squaresize}\n \\pgfmathsetmacro{\\ybottom}{-(4*(\\squaresize+\\vspace))+0.5*\\squaresize}\n\n % curly brace summarizing MSEs\n \\draw[decorate,decoration={brace,amplitude=10pt}] \n (5.8*\\squaresize,\\ytop) -- \n (5.8*\\squaresize,\\ybottom)\n node[midway,xshift=2.0cm]{$\\displaystyle \\frac{1}{5}\\sum_{i=1}^{5}\\mathrm{MSE}_i$};\n\\end{tikzpicture}\n\\end{document}\n```\n\nNow the brace endpoints are precomputed as `\\ytop` and `\\ybottom`, so TikZ has no trouble parsing the coordinates. \n\nWould you like me to also shift the brace so it hugs closer to the `MSE_i` labels, instead of floating further out?"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"status": "finished_successfully",
|
||||||
|
"end_turn": true,
|
||||||
|
"weight": 1.0,
|
||||||
|
"metadata": {
|
||||||
|
"finish_details": { "type": "stop", "stop_tokens": [200002] },
|
||||||
|
"is_complete": true,
|
||||||
|
"citations": [],
|
||||||
|
"content_references": [],
|
||||||
|
"request_id": "9830cab47b4db518-MXP",
|
||||||
|
"message_type": "next",
|
||||||
|
"model_slug": "gpt-5",
|
||||||
|
"default_model_slug": "auto",
|
||||||
|
"parent_id": "e1f63acf-2ae2-4bca-9b2b-508b705c07c0",
|
||||||
|
"timestamp_": "absolute"
|
||||||
|
},
|
||||||
|
"recipient": "all",
|
||||||
|
"channel": null
|
||||||
|
},
|
||||||
|
"parent": "e1f63acf-2ae2-4bca-9b2b-508b705c07c0",
|
||||||
|
"children": ["b357d504-4093-40bc-98cd-d2f60805af13"]
|
||||||
|
},
|
||||||
|
"b357d504-4093-40bc-98cd-d2f60805af13": {
|
||||||
|
"id": "b357d504-4093-40bc-98cd-d2f60805af13",
|
||||||
|
"message": {
|
||||||
|
"id": "b357d504-4093-40bc-98cd-d2f60805af13",
|
||||||
|
"author": { "role": "user", "name": null, "metadata": {} },
|
||||||
|
"create_time": 1758533654.699,
|
||||||
|
"update_time": null,
|
||||||
|
"content": {
|
||||||
|
"content_type": "text",
|
||||||
|
"parts": [
|
||||||
|
"Illustration of k-fold cross-validation with $k=5$. The dataset is divided into 5 folds, and the model is trained on 4 folds and evaluated on the remaining fold. This process is repeated 5 times, with each fold being used as the test set once.\n\nMake this caption better fit the illustration"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"status": "finished_successfully",
|
||||||
|
"end_turn": null,
|
||||||
|
"weight": 1.0,
|
||||||
|
"metadata": {
|
||||||
|
"selected_github_repos": [],
|
||||||
|
"serialization_metadata": { "custom_symbol_offsets": [] },
|
||||||
|
"request_id": "9830ce2f5b41b518-MXP",
|
||||||
|
"message_source": null,
|
||||||
|
"timestamp_": "absolute"
|
||||||
|
},
|
||||||
|
"recipient": "all",
|
||||||
|
"channel": null
|
||||||
|
},
|
||||||
|
"parent": "f88fa5f2-0361-4186-a8ff-cfcc7f3328ac",
|
||||||
|
"children": ["95202b75-de2d-45c5-bf51-0e6880deff19"]
|
||||||
|
},
|
||||||
|
"95202b75-de2d-45c5-bf51-0e6880deff19": {
|
||||||
|
"id": "95202b75-de2d-45c5-bf51-0e6880deff19",
|
||||||
|
"message": {
|
||||||
|
"id": "95202b75-de2d-45c5-bf51-0e6880deff19",
|
||||||
|
"author": { "role": "assistant", "name": null, "metadata": {} },
|
||||||
|
"create_time": 1758533655.141366,
|
||||||
|
"update_time": null,
|
||||||
|
"content": {
|
||||||
|
"content_type": "text",
|
||||||
|
"parts": [
|
||||||
|
"How about something like this, tailored to the way your illustration actually looks (colored test fold moving across rows, MSEs at the side, brace showing the average): \n\n**\u201cIllustration of 5-fold cross-validation. Each row corresponds to one fold: the turquoise block marks the held-out test set while the white blocks form the training set. The model\u2019s error $\\mathrm{MSE}_i$ is computed for each fold, and the final performance is the average of all $\\mathrm{MSE}_i$.\u201d**\n\nThis way the caption explicitly ties the turquoise square, the row structure, the $\\mathrm{MSE}_i$ labels, and the brace into one coherent story. \n\nWant me to also tighten it into a single-sentence version for use under a figure in a paper?"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"status": "finished_successfully",
|
||||||
|
"end_turn": true,
|
||||||
|
"weight": 1.0,
|
||||||
|
"metadata": {
|
||||||
|
"finish_details": { "type": "stop", "stop_tokens": [200002] },
|
||||||
|
"is_complete": true,
|
||||||
|
"citations": [],
|
||||||
|
"content_references": [],
|
||||||
|
"request_id": "9830ce2f5b41b518-MXP",
|
||||||
|
"message_type": "next",
|
||||||
|
"model_slug": "gpt-5",
|
||||||
|
"default_model_slug": "auto",
|
||||||
|
"parent_id": "b357d504-4093-40bc-98cd-d2f60805af13",
|
||||||
|
"timestamp_": "absolute"
|
||||||
|
},
|
||||||
|
"recipient": "all",
|
||||||
|
"channel": null
|
||||||
|
},
|
||||||
|
"parent": "b357d504-4093-40bc-98cd-d2f60805af13",
|
||||||
|
"children": []
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"moderation_results": [],
|
||||||
|
"current_node": "95202b75-de2d-45c5-bf51-0e6880deff19",
|
||||||
|
"plugin_ids": null,
|
||||||
|
"conversation_id": "68d11545-da50-8322-b59d-73ed39af2bcd",
|
||||||
|
"conversation_template_id": null,
|
||||||
|
"gizmo_id": null,
|
||||||
|
"gizmo_type": null,
|
||||||
|
"is_archived": false,
|
||||||
|
"is_starred": null,
|
||||||
|
"safe_urls": [],
|
||||||
|
"blocked_urls": [],
|
||||||
|
"default_model_slug": "auto",
|
||||||
|
"conversation_origin": null,
|
||||||
|
"voice": null,
|
||||||
|
"async_status": null,
|
||||||
|
"disabled_tool_ids": [],
|
||||||
|
"is_do_not_remember": false,
|
||||||
|
"memory_scope": "global_enabled",
|
||||||
|
"context_scopes": null,
|
||||||
|
"sugar_item_id": null,
|
||||||
|
"sugar_item_visible": false,
|
||||||
|
"is_study_mode": false,
|
||||||
|
"owner": null,
|
||||||
|
"id": "68d11545-da50-8322-b59d-73ed39af2bcd"
|
||||||
|
}
|
||||||
@@ -0,0 +1,116 @@
|
|||||||
|
|
||||||
|
@article{harris_array_2020,
|
||||||
|
title = {Array programming with {NumPy}},
|
||||||
|
volume = {585},
|
||||||
|
url = {https://doi.org/10.1038/s41586-020-2649-2},
|
||||||
|
doi = {10.1038/s41586-020-2649-2},
|
||||||
|
pages = {357--362},
|
||||||
|
number = {7825},
|
||||||
|
journaltitle = {Nature},
|
||||||
|
author = {Harris, Charles R. and Millman, K. Jarrod and Walt, Stéfan J. van der and Gommers, Ralf and Virtanen, Pauli and Cournapeau, David and Wieser, Eric and Taylor, Julian and Berg, Sebastian and Smith, Nathaniel J. and Kern, Robert and Picus, Matti and Hoyer, Stephan and Kerkwijk, Marten H. van and Brett, Matthew and Haldane, Allan and Río, Jaime Fernández del and Wiebe, Mark and Peterson, Pearu and Gérard-Marchant, Pierre and Sheppard, Kevin and Reddy, Tyler and Weckesser, Warren and Abbasi, Hameer and Gohlke, Christoph and Oliphant, Travis E.},
|
||||||
|
date = {2020-09},
|
||||||
|
year = {2020},
|
||||||
|
note = {Publisher: Springer Science and Business Media {LLC}},
|
||||||
|
}
|
||||||
|
|
||||||
|
@article{hunter_matplotlib_2007,
|
||||||
|
title = {Matplotlib: A 2D graphics environment},
|
||||||
|
volume = {9},
|
||||||
|
doi = {10.1109/MCSE.2007.55},
|
||||||
|
abstract = {Matplotlib is a 2D graphics package used for Python for application development, interactive scripting, and publication-quality image generation across user interfaces and operating systems.},
|
||||||
|
pages = {90--95},
|
||||||
|
number = {3},
|
||||||
|
journaltitle = {Computing in Science \& Engineering},
|
||||||
|
author = {Hunter, J. D.},
|
||||||
|
date = {2007},
|
||||||
|
year = {2007},
|
||||||
|
note = {Publisher: {IEEE} {COMPUTER} {SOC}},
|
||||||
|
}
|
||||||
|
|
||||||
|
@article{pedregosa_scikit-learn_2011,
|
||||||
|
title = {Scikit-learn: Machine Learning in Python},
|
||||||
|
volume = {12},
|
||||||
|
pages = {2825--2830},
|
||||||
|
journaltitle = {Journal of Machine Learning Research},
|
||||||
|
author = {Pedregosa, F. and Varoquaux, G. and Gramfort, A. and Michel, V. and Thirion, B. and Grisel, O. and Blondel, M. and Prettenhofer, P. and Weiss, R. and Dubourg, V. and Vanderplas, J. and Passos, A. and Cournapeau, D. and Brucher, M. and Perrot, M. and Duchesnay, E.},
|
||||||
|
date = {2011},
|
||||||
|
year = {2011},
|
||||||
|
}
|
||||||
|
|
||||||
|
@inreference{noauthor_stochastic_2025,
|
||||||
|
title = {Stochastic gradient descent},
|
||||||
|
rights = {Creative Commons Attribution-{ShareAlike} License},
|
||||||
|
url = {https://en.wikipedia.org/w/index.php?title=Stochastic_gradient_descent&oldid=1309164477},
|
||||||
|
abstract = {Stochastic gradient descent (often abbreviated {SGD}) is an iterative method for optimizing an objective function with suitable smoothness properties (e.g. differentiable or subdifferentiable). It can be regarded as a stochastic approximation of gradient descent optimization, since it replaces the actual gradient (calculated from the entire data set) by an estimate thereof (calculated from a randomly selected subset of the data). Especially in high-dimensional optimization problems this reduces the very high computational burden, achieving faster iterations in exchange for a lower convergence rate.
|
||||||
|
The basic idea behind stochastic approximation can be traced back to the Robbins–Monro algorithm of the 1950s. Today, stochastic gradient descent has become an important optimization method in machine learning.},
|
||||||
|
booktitle = {Wikipedia},
|
||||||
|
urldate = {2025-09-22},
|
||||||
|
date = {2025-09-02},
|
||||||
|
year = {2025},
|
||||||
|
langid = {english},
|
||||||
|
note = {Page Version {ID}: 1309164477},
|
||||||
|
file = {Snapshot:/home/lars/Zotero/storage/B6CVR59B/index.html:text/html},
|
||||||
|
}
|
||||||
|
|
||||||
|
@unpublished{elstner_lecture_2025,
|
||||||
|
location = {Karlsruhe Institute for Technology, Karlsruhe},
|
||||||
|
title = {Lecture: Machine Learning for Chemistry},
|
||||||
|
type = {Lecture},
|
||||||
|
howpublished = {Lecture},
|
||||||
|
author = {Elstner, Marcus and Kubar, Tomas},
|
||||||
|
date = {2025-05-20},
|
||||||
|
year = {2025},
|
||||||
|
langid = {german},
|
||||||
|
file = {PDF:/home/lars/Zotero/storage/5LSLJMK8/Elstner and Kubar - 2025 - Lecture Machine Learning for Chemistry.pdf:application/pdf},
|
||||||
|
}
|
||||||
|
|
||||||
|
@online{lekhansh_lasso_2024,
|
||||||
|
title = {Lasso vs. Ridge Regression: A Detailed Comparison},
|
||||||
|
url = {https://medium.com/@tyagi.lekhansh/lasso-vs-ridge-regression-a-detailed-comparison-140f7832c624},
|
||||||
|
shorttitle = {Lasso vs. Ridge Regression},
|
||||||
|
abstract = {In the realm of regression analysis, Lasso (Least Absolute Shrinkage and Selection Operator) and Ridge Regression are two popular…},
|
||||||
|
titleaddon = {Medium},
|
||||||
|
author = {Lekhansh},
|
||||||
|
urldate = {2025-09-22},
|
||||||
|
date = {2024-09-04},
|
||||||
|
year = {2024},
|
||||||
|
langid = {english},
|
||||||
|
}
|
||||||
|
|
||||||
|
@book{hastie_elements_2009,
|
||||||
|
location = {New York, {NY}},
|
||||||
|
title = {The Elements of Statistical Learning},
|
||||||
|
rights = {http://www.springer.com/tdm},
|
||||||
|
isbn = {978-0-387-84857-0 978-0-387-84858-7},
|
||||||
|
url = {http://link.springer.com/10.1007/978-0-387-84858-7},
|
||||||
|
series = {Springer Series in Statistics},
|
||||||
|
publisher = {Springer},
|
||||||
|
author = {Hastie, Trevor and Tibshirani, Robert and Friedman, Jerome},
|
||||||
|
urldate = {2025-09-22},
|
||||||
|
date = {2009},
|
||||||
|
year = {2009},
|
||||||
|
doi = {10.1007/978-0-387-84858-7},
|
||||||
|
keywords = {Averaging, Boosting, classification, clustering, data mining, machine learning, Projection pursuit, Random Forest, supervised learning, Support Vector Machine, unsupervised learning},
|
||||||
|
file = {Full Text PDF:/home/lars/Zotero/storage/D3N4DVY9/Hastie et al. - 2009 - The Elements of Statistical Learning.pdf:application/pdf},
|
||||||
|
}
|
||||||
|
|
||||||
|
@book{goodfellow_deep_2016,
|
||||||
|
title = {Deep Learning},
|
||||||
|
publisher = {{MIT} Press},
|
||||||
|
author = {Goodfellow, Ian and Bengio, Yoshua and Courville, Aaron},
|
||||||
|
date = {2016},
|
||||||
|
year = {2016},
|
||||||
|
}
|
||||||
|
|
||||||
|
@book{bishop_pattern_2006,
|
||||||
|
location = {New York},
|
||||||
|
title = {Pattern recognition and machine learning},
|
||||||
|
isbn = {978-0-387-31073-2},
|
||||||
|
series = {Information science and statistics},
|
||||||
|
publisher = {Springer},
|
||||||
|
author = {Bishop, Christopher M.},
|
||||||
|
date = {2006},
|
||||||
|
year = {2006},
|
||||||
|
langid = {english},
|
||||||
|
file = {PDF:/home/lars/Zotero/storage/9H5W9BGC/Bishop - 2006 - Pattern recognition and machine learning.pdf:application/pdf},
|
||||||
|
}
|
||||||
|
|||||||
Binary file not shown.
+7
-3
@@ -18,6 +18,10 @@
|
|||||||
\usepackage{float}
|
\usepackage{float}
|
||||||
\usepackage{todonotes}
|
\usepackage{todonotes}
|
||||||
|
|
||||||
|
% tikz decorations
|
||||||
|
\usepackage{tikz}
|
||||||
|
\usetikzlibrary{decorations.pathreplacing}
|
||||||
|
|
||||||
\usepackage{upgreek} %upalpha in Saxena2021 Reference
|
\usepackage{upgreek} %upalpha in Saxena2021 Reference
|
||||||
\usepackage[utf8]{inputenc}
|
\usepackage[utf8]{inputenc}
|
||||||
\usepackage{hyperref}
|
\usepackage{hyperref}
|
||||||
@@ -29,6 +33,7 @@
|
|||||||
|
|
||||||
|
|
||||||
\usepackage{xcolor}
|
\usepackage{xcolor}
|
||||||
|
\definecolor{turq}{RGB}{64,224,208}
|
||||||
%\newcommand{\contrib}[1]{\textcolor{red}{#1}}
|
%\newcommand{\contrib}[1]{\textcolor{red}{#1}}
|
||||||
%\newcommand{\comment}[1]{\textcolor{blue}{#1}}
|
%\newcommand{\comment}[1]{\textcolor{blue}{#1}}
|
||||||
|
|
||||||
@@ -75,8 +80,7 @@ apsrev41Control%
|
|||||||
|
|
||||||
|
|
||||||
\begin{abstract}
|
\begin{abstract}
|
||||||
Advances in artificial intelligence/machine learning methods provide tools that have broad applicability in scientific research. These techniques are being applied across the diversity of nuclear physics research topics, leading to advances that will facilitate scientific discoveries and societal applications.
|
The basis for many data-driven methods is the optimization of numerical models to describe relations in data. With the increasing availability of large datasets and the growing complexity of models, there is an increasing need for efficient optimization techniques that can handle the computational challenges associated with big data. Furthermore, a good optimization metric can be crucial to find the best possible model to describe the data. In this paper different methods of such optimization are studied in the context of large datasets and big data with possibly noise-affected input data. To render the best possible fits a multitude of cost functions, minimization algorithms and other approaches to reduce the computation associated with the optimization process are studied and evaluated. Reducing the computational cost is also important to improve the economical and ecological footprint of training large models on big data, as the spread of datadriven methods in all areas of life is increasing rapidly.
|
||||||
This review provides a snapshot of nuclear physics research which has been transformed by artificial intelligence and machine learning techniques.
|
|
||||||
\end{abstract}
|
\end{abstract}
|
||||||
|
|
||||||
\maketitle
|
\maketitle
|
||||||
@@ -91,7 +95,7 @@ apsrev41Control%
|
|||||||
\section{Methods} \label{sec:methods}
|
\section{Methods} \label{sec:methods}
|
||||||
\input{chapters/methods.tex}
|
\input{chapters/methods.tex}
|
||||||
|
|
||||||
\section{Results} \label{sec:results}
|
\section{Results and Discussion} \label{sec:results}
|
||||||
\input{chapters/results.tex}
|
\input{chapters/results.tex}
|
||||||
|
|
||||||
\section{Conclusion} \label{sec:conclusion}
|
\section{Conclusion} \label{sec:conclusion}
|
||||||
|
|||||||
+29
-14
@@ -50,7 +50,9 @@
|
|||||||
"source": [
|
"source": [
|
||||||
"x = np.linspace(-1, 1, 100_000)\n",
|
"x = np.linspace(-1, 1, 100_000)\n",
|
||||||
"y = datamanip.noise_data(datamanip.runge_function(x), 1.0)\n",
|
"y = datamanip.noise_data(datamanip.runge_function(x), 1.0)\n",
|
||||||
"x_train, x_test, y_train, y_test = train_test_split(x, y, test_size=0.2)\n",
|
"x_train, x_test, y_train, y_test = train_test_split(\n",
|
||||||
|
" x, y, test_size=0.2, random_state=datamanip.get_RNG().integers(0, 1e6)\n",
|
||||||
|
")\n",
|
||||||
"\n",
|
"\n",
|
||||||
"fig, ax = plotting.scatter_dataset(x_train, x_test, y_train, y_test)\n",
|
"fig, ax = plotting.scatter_dataset(x_train, x_test, y_train, y_test)\n",
|
||||||
"fig.set_layout_engine(\"compressed\")\n",
|
"fig.set_layout_engine(\"compressed\")\n",
|
||||||
@@ -255,6 +257,7 @@
|
|||||||
" [f\"${format_number(tick)}$\" for tick in lambda_values[::3]], rotation=45\n",
|
" [f\"${format_number(tick)}$\" for tick in lambda_values[::3]], rotation=45\n",
|
||||||
")\n",
|
")\n",
|
||||||
"\n",
|
"\n",
|
||||||
|
"fig.tight_layout()\n",
|
||||||
"fig.savefig(os.path.join(FIG_DIR, \"ridge_parameter_plot.pdf\"))"
|
"fig.savefig(os.path.join(FIG_DIR, \"ridge_parameter_plot.pdf\"))"
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
@@ -477,7 +480,9 @@
|
|||||||
"\n",
|
"\n",
|
||||||
"x = np.linspace(-1, 1, X_size)\n",
|
"x = np.linspace(-1, 1, X_size)\n",
|
||||||
"y = datamanip.noise_data(datamanip.runge_function(x), 0.1) # LOWER NOISE FOR SGD\n",
|
"y = datamanip.noise_data(datamanip.runge_function(x), 0.1) # LOWER NOISE FOR SGD\n",
|
||||||
"x_train, x_test, y_train, y_test = train_test_split(x, y, test_size=0.2)\n",
|
"x_train, x_test, y_train, y_test = train_test_split(\n",
|
||||||
|
" x, y, test_size=0.2, random_state=datamanip.get_RNG().integers(0, 1e6)\n",
|
||||||
|
")\n",
|
||||||
"\n",
|
"\n",
|
||||||
"X_train = datamanip.polynomial_features(x_train, 10, False)\n",
|
"X_train = datamanip.polynomial_features(x_train, 10, False)\n",
|
||||||
"X_test = datamanip.polynomial_features(x_test, 10, False)\n",
|
"X_test = datamanip.polynomial_features(x_test, 10, False)\n",
|
||||||
@@ -518,7 +523,9 @@
|
|||||||
"\n",
|
"\n",
|
||||||
"x = np.linspace(-1, 1, X_size)\n",
|
"x = np.linspace(-1, 1, X_size)\n",
|
||||||
"y = datamanip.noise_data(datamanip.runge_function(x), 0.1) # LOWER NOISE FOR SGD\n",
|
"y = datamanip.noise_data(datamanip.runge_function(x), 0.1) # LOWER NOISE FOR SGD\n",
|
||||||
"x_train, x_test, y_train, y_test = train_test_split(x, y, test_size=0.2)\n",
|
"x_train, x_test, y_train, y_test = train_test_split(\n",
|
||||||
|
" x, y, test_size=0.2, random_state=datamanip.get_RNG().integers(0, 1e6)\n",
|
||||||
|
")\n",
|
||||||
"\n",
|
"\n",
|
||||||
"X_train = datamanip.polynomial_features(x_train, 10, False)\n",
|
"X_train = datamanip.polynomial_features(x_train, 10, False)\n",
|
||||||
"X_test = datamanip.polynomial_features(x_test, 10, False)\n",
|
"X_test = datamanip.polynomial_features(x_test, 10, False)\n",
|
||||||
@@ -603,14 +610,16 @@
|
|||||||
"metadata": {},
|
"metadata": {},
|
||||||
"outputs": [],
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"x = np.linspace(-1, 1, 300)\n",
|
"x = np.linspace(-1, 1, 100)\n",
|
||||||
"y = datamanip.runge_function(x)\n",
|
"y = datamanip.runge_function(x)\n",
|
||||||
"x_train, x_test, y_train, y_test_noise_free = train_test_split(x, y, test_size=0.2)\n",
|
"x_train, x_test, y_train, y_test_noise_free = train_test_split(\n",
|
||||||
|
" x, y, test_size=0.2, random_state=datamanip.get_RNG().integers(0, 1e6)\n",
|
||||||
|
")\n",
|
||||||
"y_train = datamanip.noise_data(y_train, 1.0)\n",
|
"y_train = datamanip.noise_data(y_train, 1.0)\n",
|
||||||
"y_test = datamanip.noise_data(y_test_noise_free, 1.0)\n",
|
"y_test = datamanip.noise_data(y_test_noise_free, 1.0)\n",
|
||||||
"\n",
|
"\n",
|
||||||
"polynomial_degrees = np.arange(1, 50)\n",
|
"polynomial_degrees = np.arange(1, 15)\n",
|
||||||
"n_bootstraps = len(x_train)\n",
|
"n_bootstraps = 5000\n",
|
||||||
"mses = np.zeros((len(polynomial_degrees), n_bootstraps))\n",
|
"mses = np.zeros((len(polynomial_degrees), n_bootstraps))\n",
|
||||||
"biases = np.zeros((len(polynomial_degrees), n_bootstraps))\n",
|
"biases = np.zeros((len(polynomial_degrees), n_bootstraps))\n",
|
||||||
"variances = np.zeros((len(polynomial_degrees), n_bootstraps))\n",
|
"variances = np.zeros((len(polynomial_degrees), n_bootstraps))\n",
|
||||||
@@ -719,9 +728,14 @@
|
|||||||
"metadata": {},
|
"metadata": {},
|
||||||
"outputs": [],
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"fig, (ax, ax2) = plt.subplots(\n",
|
"SINGULAR_PLOT = True\n",
|
||||||
" 1, 2, figsize=plotting.get_figsize(0.35, True), sharey=True\n",
|
"if not SINGULAR_PLOT:\n",
|
||||||
")\n",
|
" fig, (ax, ax2) = plt.subplots(\n",
|
||||||
|
" 1, 2, figsize=plotting.get_figsize(0.35, True), sharey=True\n",
|
||||||
|
" )\n",
|
||||||
|
"else:\n",
|
||||||
|
" fig, ax = plt.subplots(1, 1, figsize=plotting.get_figsize(0.5, False))\n",
|
||||||
|
" ax2 = ax\n",
|
||||||
"mse_mean = np.mean(mses, axis=1)\n",
|
"mse_mean = np.mean(mses, axis=1)\n",
|
||||||
"bias_mean = np.mean(biases, axis=1)\n",
|
"bias_mean = np.mean(biases, axis=1)\n",
|
||||||
"var_mean = np.mean(variances, axis=1)\n",
|
"var_mean = np.mean(variances, axis=1)\n",
|
||||||
@@ -767,13 +781,14 @@
|
|||||||
"\n",
|
"\n",
|
||||||
"fig.tight_layout()\n",
|
"fig.tight_layout()\n",
|
||||||
"\n",
|
"\n",
|
||||||
"\n",
|
"if not SINGULAR_PLOT:\n",
|
||||||
"ax2.plot(polynomial_degrees, mse_mean, label=\"Bootstrapping MSE\", color=\"C0\")\n",
|
" ax2.plot(polynomial_degrees, mse_mean, label=\"Bootstrapping MSE\", color=\"C1\")\n",
|
||||||
"ax2.plot(\n",
|
"ax2.plot(\n",
|
||||||
" polynomial_degrees,\n",
|
" polynomial_degrees,\n",
|
||||||
" np.mean(k_fold_mses, axis=1),\n",
|
" np.mean(k_fold_mses, axis=1),\n",
|
||||||
" label=f\"{k_folds}-Fold Crossvalidation MSE\",\n",
|
" label=f\"{k_folds}-Fold Crossvalidation MSE\",\n",
|
||||||
" color=\"C1\",\n",
|
" color=\"black\",\n",
|
||||||
|
" ls=\"--\",\n",
|
||||||
")\n",
|
")\n",
|
||||||
"\n",
|
"\n",
|
||||||
"\n",
|
"\n",
|
||||||
@@ -791,7 +806,7 @@
|
|||||||
"metadata": {},
|
"metadata": {},
|
||||||
"outputs": [],
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"polynomial_degrees = np.arange(1, 30)\n",
|
"polynomial_degrees = np.arange(1, 25)\n",
|
||||||
"k_folds = 5\n",
|
"k_folds = 5\n",
|
||||||
"\n",
|
"\n",
|
||||||
"k_fold_mses_ols = np.zeros((len(polynomial_degrees), k_folds))\n",
|
"k_fold_mses_ols = np.zeros((len(polynomial_degrees), k_folds))\n",
|
||||||
|
|||||||
@@ -5,6 +5,10 @@ from sklearn.utils import resample # type: ignore
|
|||||||
from sklearn.model_selection import KFold # type: ignore
|
from sklearn.model_selection import KFold # type: ignore
|
||||||
|
|
||||||
|
|
||||||
|
def get_RNG() -> np.random.Generator:
|
||||||
|
return np.random.default_rng(314)
|
||||||
|
|
||||||
|
|
||||||
def polynomial_features(x: np.ndarray, p: int, intercept: bool = True) -> np.ndarray:
|
def polynomial_features(x: np.ndarray, p: int, intercept: bool = True) -> np.ndarray:
|
||||||
"""Generates a design matrix with polynomial features up to degree p.
|
"""Generates a design matrix with polynomial features up to degree p.
|
||||||
Args:
|
Args:
|
||||||
@@ -69,7 +73,7 @@ def noise_data(y: np.ndarray, noise_level: float = 1.0) -> np.ndarray:
|
|||||||
Returns:
|
Returns:
|
||||||
The noisy target vector of shape (n_samples,).
|
The noisy target vector of shape (n_samples,).
|
||||||
"""
|
"""
|
||||||
noise = np.random.normal(0, noise_level, size=y.shape)
|
noise = get_RNG().normal(0, noise_level, size=y.shape)
|
||||||
return y + noise
|
return y + noise
|
||||||
|
|
||||||
|
|
||||||
@@ -97,7 +101,9 @@ def bootstrap_resample(
|
|||||||
"""
|
"""
|
||||||
resamples = []
|
resamples = []
|
||||||
for _ in range(n_resamples):
|
for _ in range(n_resamples):
|
||||||
X_resample, y_resample = resample(X, y)
|
X_resample, y_resample = resample(
|
||||||
|
X, y, random_state=get_RNG().integers(0, int(1e6))
|
||||||
|
)
|
||||||
resamples.append((X_resample, y_resample))
|
resamples.append((X_resample, y_resample))
|
||||||
return resamples
|
return resamples
|
||||||
|
|
||||||
@@ -113,7 +119,7 @@ def k_fold_split(
|
|||||||
Returns:
|
Returns:
|
||||||
A list of tuples, each containing (X_train, y_train, X_val, y_val) for each fold.
|
A list of tuples, each containing (X_train, y_train, X_val, y_val) for each fold.
|
||||||
"""
|
"""
|
||||||
kf = KFold(n_splits=k, shuffle=True)
|
kf = KFold(n_splits=k, shuffle=True, random_state=get_RNG().integers(0, int(1e6)))
|
||||||
folds = []
|
folds = []
|
||||||
for train_index, val_index in kf.split(X):
|
for train_index, val_index in kf.split(X):
|
||||||
X_train, X_val = X[train_index], X[val_index]
|
X_train, X_val = X[train_index], X[val_index]
|
||||||
|
|||||||
@@ -1,4 +1,5 @@
|
|||||||
import numpy as np
|
import numpy as np
|
||||||
|
from pyoptim.datamanip import get_RNG
|
||||||
|
|
||||||
|
|
||||||
def Ridge_parameters(X: np.ndarray, y: np.ndarray, lam: float) -> np.ndarray:
|
def Ridge_parameters(X: np.ndarray, y: np.ndarray, lam: float) -> np.ndarray:
|
||||||
@@ -274,17 +275,18 @@ class OLSStochasticGradientDescent(OLSGradientDescent):
|
|||||||
super().__init__(*args, **kwargs)
|
super().__init__(*args, **kwargs)
|
||||||
self.batch_size = batch_size
|
self.batch_size = batch_size
|
||||||
self.batches_per_epoch = batches_per_epoch
|
self.batches_per_epoch = batches_per_epoch
|
||||||
|
self.RNG = get_RNG()
|
||||||
|
|
||||||
def _precomp(self):
|
def _precomp(self):
|
||||||
self.N = len(self.y)
|
self.N = len(self.y)
|
||||||
self.indices = np.arange(self.N)
|
self.indices = np.arange(self.N)
|
||||||
self.n = self.batch_size
|
self.n = self.batch_size
|
||||||
np.random.shuffle(self.indices)
|
self.RNG.shuffle(self.indices)
|
||||||
self.X = self.X[self.indices]
|
self.X = self.X[self.indices]
|
||||||
self.y = self.y[self.indices]
|
self.y = self.y[self.indices]
|
||||||
|
|
||||||
def _comp_step(self):
|
def _comp_step(self):
|
||||||
index = np.random.randint(0, self.N)
|
index = self.RNG.integers(0, self.N)
|
||||||
batch_indices = slice(index, index + self.batch_size)
|
batch_indices = slice(index, index + self.batch_size)
|
||||||
if index + self.batch_size > self.N:
|
if index + self.batch_size > self.N:
|
||||||
batch_indices = slice(index, self.N)
|
batch_indices = slice(index, self.N)
|
||||||
|
|||||||
Reference in New Issue
Block a user