Covariance of observed and fitted values
Covariance of observed and fitted values
Loading saved threads...
Makas · External communityPost link
External question — Cross Validated Stack Exchange
Author: Makas
Original post: https://stats.stackexchange.com/questions/662681
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
I am confused about several computations I've seen for the covariance between the response and fitted values in linear regression.
For instance, it is a standard step to derive the bias-variance trade-off to show that
$$
\mathbb{E}[(Y - f(x_0))(\hat f(x_0) - \mathbb{E}\hat f(x_0)) | X = x_0] = 0
$$
see e.g. Equation (7.9) in
Elements of Statistical Learning
.
My understanding is that the key step in this computation is that, if we first condition on the training data set
$\mathcal T = \{(x_1,y_1),\dots, (x_n,y_n)\}$
, then
$\hat f(x_0) - \mathbb{E}\hat f(x_0)$
is constant for a fixed
$X=x_0$
and thus
$$
\begin{align}\tag{1}\label{eq:1}
\mathbb{E}[(Y - f(x_0))(\hat f(x_0) - \mathbb{E}\hat f(x_0)) | X = x_0,\mathcal T] &= (\hat f(x_0) - \mathbb{E}\hat f(x_0))\mathbb{E}[(Y - f(x_0)) | X = x_0,\mathcal T] \\ &= 0.
\end{align}
$$
On the other hand, I've seen computations of
$$
\operatorname{Cov}(Y_i, \hat Y_i) \neq 0,
$$
see e.g. the first answer in
this post
.
To be precise, I think these computations deal with
$\operatorname{Cov}(Y_i, \hat Y_i) = \operatorname{Cov}(Y,\hat f(x_i) | X = x_i) = \mathbb{E}[(Y - f(x_i))(\hat f(x_i) - \mathbb{E}\hat f(x_i)) | X = x_i]$
, where now
$x_i$
is not arbitrary but instead is a point in the training set.
In any case, Equation \eqref{eq:1} should be valid for any
$X = x_0$
, in particular it should also be valid for
$X = x_i$
a point in the training set.
I'd really appreciate if someone could explain how these two apparently contradicting results are compatible.
I have a feeling that it is all about what one is conditioning on, but it seems to me like both cases condition only on a single input point
$X = x_0$
and average over both
$Y = f(x_0) + \epsilon$
(randomness in
$\epsilon$
) as well as
$\hat f(x_0) = \hat f(x_0;\mathbf{y})$
(randomness in
$\mathbf{y} = (y_1,\dots,y_n)$
; note that
$\hat f(x_0)$
depends on the training data through
$\hat f(x_0) = x_0^T X\hat\beta = x_0^T X(X^TX)^{-1}X^T\mathbf{y}$
.
Let's assume that
$X$
is constant for simplicity.)
Quote
Report
Post Reply
Quoted from Forex.com.bd-Editorial External question — Cross Validated Stack Exchange Author: Makas Source score (net votes, not local likes): 0 Original post: https://stats.stackexchange.com/questions/662681 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. I am confused about several computations I've seen for the covariance between the response and fitted values in linear regression. For instance, it is a standard step to derive the bias-variance trade-off to show that $$ \mathbb{E}[(Y - f(x_0))(\hat f(x_0) - \mathbb{E}\hat f(x_0)) | X = x_0] = 0 $$ see e.g. Equation (7.9) in Elements of Statistical Learning . My understanding is that the key step in this computation is that, if we first condition on the training data set $\mathcal T = \{(x_1,y_1),\dots, (x_n,y_n)\}$ , then $\hat f(x_0) - \mathbb{E}\hat f(x_0)$ is constant for a fixed $X=x_0$ and thus $$ \begin{align}\tag{1}\label{eq:1} \mathbb{E}[(Y - f(x_0))(\hat f(x_0) - \mathbb{E}\hat f(x_0)) | X = x_0,\mathcal T] &= (\hat f(x_0) - \mathbb{E}\hat f(x_0))\mathbb{E}[(Y - f(x_0)) | X = x_0,\mathcal T] \\ &= 0. \end{align} $$ On the other hand, I've seen computations of $$ \operatorname{Cov}(Y_i, \hat Y_i) \neq 0, $$ see e.g. the first answer in this post . To be precise, I think these computations deal with $\operatorname{Cov}(Y_i, \hat Y_i) = \operatorname{Cov}(Y,\hat f(x_i) | X = x_i) = \mathbb{E}[(Y - f(x_i))(\hat f(x_i) - \mathbb{E}\hat f(x_i)) | X = x_i]$ , where now $x_i$ is not arbitrary but instead is a point in the training set. In any case, Equation \eqref{eq:1} should be valid for any $X = x_0$ , in particular it should also be valid for $X = x_i$ a point in the training set. I'd really appreciate if someone could explain how these two apparently contradicting results are compatible. I have a feeling that it is all about what one is conditioning on, but it seems to me like both cases condition only on a single input point $X = x_0$ and average over both $Y = f(x_0) + \epsilon$ (randomness in $\epsilon$ ) as well as $\hat f(x_0) = \hat f(x_0;\mathbf{y})$ (randomness in $\mathbf{y} = (y_1,\dots,y_n)$ ; note that $\hat f(x_0)$ depends on the training data through $\hat f(x_0) = x_0^T X\hat\beta = x_0^T X(X^TX)^{-1}X^T\mathbf{y}$ . Let's assume that $X$ is constant for simplicity.)
Checking account access…