Poor performance of least squares when $\mathbf{X}'\mathbf{X}$ is "not nearly a unit correlation matrix", Hoerl and Kennard (1970)
Poor performance of least squares when $\mathbf{X}'\mathbf{X}$ is "not nearly a unit correlation matrix", Hoerl and Kennard (1970)
Loading saved threads...
microhaus · External communityPost link
External question — Cross Validated Stack Exchange
Author: microhaus
Original post: https://stats.stackexchange.com/questions/661240
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
Context.
In
Ridge Regression: Biased Estimation of Nonorthogonal Problems
, by Hoerl and Kennard (1970), the authors repeatedly use the following observation to motivate analytic study of both the behaviour and properties of the ridge regression estimator. From
Section 0: Introduction
,
The usual estimation procedure for the unknown
$\boldsymbol{\beta}$
is Gauss-Markov --linear functions of
$\mathbf{Y} = \{y_v\}$
that are unbiased and have minimum variance. This estimation procedure is a good one if
$\mathbf{X}'\mathbf{X}$
, when in the form of a correlation matrix, is nearly a unit matrix. However, if
$\mathbf{X}'\mathbf{X}$
is not nearly a unit matrix, the least squares estimates are sensitive to a number of "errors". THe results of these errors are critical when the specification is that
$\mathbf{X} \boldsymbol{\beta}$
is a true model.
The same phrasing of "
$(\mathbf{X}'\mathbf{X})$
, when written as a correlation matrix, is not nearly a unit matrix" can further be found in
Section 1: Properties of Brst Linear Unbiased Estimation
and
Section 3a: Definition of the Ridge Trace
.
In almost all the modern presentations of the ridge regression estimator I've read, such as in Hastie et al. (2010), none of them tend to be couched in the language of "unit correlation matrices".
Instead, the modern presentation tends to motivate ridge regression estimators as trading off bias for decreased variance. A little more specifically, by analysing when the variance of the least squares estimator
$\sigma^2 (\mathbf{X}'\mathbf{X})^{-1}$
can "blow up" via eigendecomposition or singular value analysis of
$(\mathbf{X}'\mathbf{X})^{-1}$
.
Questions.
1. Please may someone clarify with an argument what exactly is meant here?
2. Has this unit correlation matrix argument got anything to do with the fact that we assume the following about the errors $\mathbb{E}[\mathbf{e}\mathbf{e}'] = \sigma^2I$, i.e. an isotropic covariance matrix $\sigma^2I$ on $\mathbf{Y}$?$^{[1]}$
$^{[1]}$
In the sense that the least squares estimator is not robust to misspecification of our modelling assumption that errors have isotropic covariance?*
Related questions.
This was also asked on MathStackexchange:
https://math.stackexchange.com/questions/2340791/ridge-regression-unit-matrix-hoerl-and-kennard-1970
Quote
Report
Michael Hardy · External communityPost link
External answer — Cross Validated Stack Exchange
Author: Michael Hardy
Original post: https://stats.stackexchange.com/a/661252
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
To say that
$\mathbf X'\mathbf X$
is "nearly a unit matrix" in this case would mean that columns of the design matrix
$\mathbf X$
are nearly uncorrelated with each other.
Imagine, for example, that you have data indicating how much income each of
$n$
persons had last year and how much they spent last year. Probably those are highly correlated with each other. Now suppose you want to regress some variable related to a person's health on those two predictors. You get least-squares estimates for both predictors. But how since these two predictors highly correlated, it is hard to tell how much of the variation in the response variable is explained by one of them and how much by the other. Maybe you could drastically reduce the coefficient of one of them and drastically increase that of the other and get a sum of squares of residuals only microscopically bigger than what you got with least squares.
Quote
Report
Ben · External communityPost link
External answer — Cross Validated Stack Exchange
Author: Ben
Original post: https://stats.stackexchange.com/a/661260
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
To understand the connections between these various concepts and issues, I think you would find it very helpful to have a look at some material on geometric interpretations of linear regression and correlation.
O'Neill (2019)
) gives a broad analysis of how linear regression can be framed in geometric terms with respect to various correlation matrices, etc., and how these various concepts relate to each other. Much of what I will say here comes from that paper.
Framing the regression in terms of correlation matrices and related geometric objects:
To begin with, the matrix
$\mathbf{x}^\text{T}\mathbf{x}$
is
not
a correlation matrix --- it is the
Gramian matrix
for the regression. Suppose we define the
design correlation matrix
(containing the sample correlations for the columns of the design matrix) by:
$$\boldsymbol{\Omega}
\equiv \begin{bmatrix}
\text{Corr}(\mathbf{y},\mathbf{x}_1) \\[6pt]
\text{Corr}(\mathbf{y},\mathbf{x}_2) \\[6pt]
\vdots \\[6pt]
\text{Corr}(\mathbf{y},\mathbf{x}_m) \\[6pt]
\end{bmatrix}
\quad \quad \quad
\boldsymbol{\Theta}
\equiv \begin{bmatrix}
\text{Corr}(\mathbf{x}_1,\mathbf{x}_1) & \text{Corr}(\mathbf{x}_1,\mathbf{x}_2) & \cdots & \text{Corr}(\mathbf{x}_1,\mathbf{x}_m) \\[6pt]
\text{Corr}(\mathbf{x}_2,\mathbf{x}_1) & \text{Corr}(\mathbf{x}_2,\mathbf{x}_2) & \cdots & \text{Corr}(\mathbf{x}_2,\mathbf{x}_m) \\[6pt]
\vdots & \vdots & \ddots & \vdots \\[6pt]
\text{Corr}(\mathbf{x}_m,\mathbf{x}_1) & \text{Corr}(\mathbf{x}_m,\mathbf{x}_2) & \cdots & \text{Corr}(\mathbf{x}_m,\mathbf{x}_m) \\[6pt]
\end{bmatrix}.$$
We can also define the matrix of lengths:
$$||\mathbf{y}||
\equiv \begin{bmatrix}
||\mathbf{y}_1|| \\[6pt]
||\mathbf{y}_2|| \\[6pt]
\vdots \\[6pt]
||\mathbf{y}_m|| \\[6pt]
\end{bmatrix}
\quad \quad \quad \quad \quad
||\mathbf{x}||
\equiv \begin{bmatrix}
||\mathbf{x}_1|| & 0 & \cdots & 0 \\[6pt]
0 & ||\mathbf{x}_2|| & \ddots & 0 \\[6pt]
\vdots & \vdots & \ddots & \vdots \\[6pt]
0 & 0 & \cdots & ||\mathbf{x}_m|| \\[6pt]
\end{bmatrix}.$$
We then have
$\mathbb{x}^\text{T} \mathbb{x} = ||\mathbf{x}|| \boldsymbol{\Theta} ||\mathbf{x}||$
and
$\mathbb{x}^\text{T} \mathbb{y} = ||\mathbf{y}|| ||\mathbf{x}|| \boldsymbol{\Omega}$
, which gives the coefficient of determination for the regression as
$R^2 = \boldsymbol{\Omega}^\text{T} \boldsymbol{\Theta}^{-1} \boldsymbol{\Omega}$
.
The variance of the OLS estimator:
Taking
$\boldsymbol{\Theta} = \mathbf{v} \boldsymbol{\Lambda} \mathbf{v}^\text{T}$
to be the spectral decomposition of the design correlation matrix, and taking
$\mathbf{r} \equiv (\mathbf{v} ||\mathbf{x}||)^{-1}$
the variance of the OLS estimator is:
$$\begin{align}
\mathbb{V}(\hat{\boldsymbol{\beta}})
&= \sigma^2 (\mathbf{x}^\text{T} \mathbf{x})^{-1} \\[6pt]
&= \sigma^2 ||\mathbf{x}||^{-1} \boldsymbol{\Theta}^{-1} ||\mathbf{x}||^{-1} \\[6pt]
&= \sigma^2 ||\mathbf{x}||^{-1} (\mathbf{v} \boldsymbol{\Lambda} \mathbf{v}^\text{T})^{-1} ||\mathbf{x}||^{-1} \\[6pt]
&= \sigma^2 ||\mathbf{x}||^{-1} \mathbf{v}^{-1} \boldsymbol{\Lambda}^{-1} (\mathbf{v}^\text{T})^{-1} ||\mathbf{x}||^{-1} \\[6pt]
&= \sigma^2 (\mathbf{v} ||\mathbf{x}||)^{-1} \boldsymbol{\Lambda}^{-1} (||\mathbf{x}|| \mathbf{v}^\text{T})^{-1} \\[6pt]
&= \sigma^2 \mathbf{r} \boldsymbol{\Lambda}^{-1} \mathbf{r}^\text{T}. \\[6pt]
\end{align}$$
As you can see, this variance can be framed in a number of ways. It can be framed directly in terms of the inverse Gramian matrix
$(\mathbf{x}^\text{T} \mathbf{x})^\text{-1}$
, or in terms of the inversion of the design correlation matrix and associated length matrices, or in terms of the inversion of the underlying spectral decomposition of the design correlation matrix and related principal components. In all cases, the inverted term in the middle can "blow up" if the non-inverted term is far from proportionality to the identity matrix. It is legitimate to examine the variance of the OLS estimator based on any of these "three levels" of analysis, depending on what input is of particular interest.
As to your secondary question, yes, these results hinge on the assumed error variance
$\mathbb{V}(\boldsymbol{\varepsilon}|\mathbf{x}) = \sigma^2 \mathbf{I}$
. If the error variance is changed (e.g., to incorporate heteroscedasticity and/or correlated errors) then this adds a weighting term to the variance of the OLS estimator and the above results change accordingly. As mentioned in
this related answer
, autocorrelated errors do not cause bias in the OLS estimator but they do affect the variance of the estimator.
Quote
Report
Post Reply
Checking account access…