What we want

We are interested in the generalization error on all data drawn from the true model, not just the error on the particular training dataset \( \mathcal{L} \) that we have in hand. This is just the expectation of the cost function over many different data sets \( \{\mathcal{L}_j\} \). Denote this expectation value by \( E_{\mathcal{L}} \). In other words, we can view \( \hat{g}_{\mathcal{L}} \) as a stochastic functional that depends on the dataset \( \mathcal{L} \) and we can think of \( E_{\mathcal{L}} \) as the expected value of the functional if we drew an infinite number of datasets \( \{\mathcal{L}_1, \mathcal{L}_2, \ldots \} \).