Regression with noises in X. Should I use the unbiased estimator or the OLS estimator for forecasting?

Regression with noises in X. Should I use the unbiased estimator or the OLS estimator for forecasting?

Manage alerts

Loading saved threads...

The One · External communityPost link
External question — Cross Validated Stack Exchange Author: The One Original post: https://stats.stackexchange.com/questions/650835 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. I am working with a dataset that includes variables $Y$ and $X$ . I assume that $$ Y = \beta X + \epsilon $$ satisfies all the assumptions of OLS. Based on industry knowledge, I know that theoretically $\beta = 1$ . However, I am aware that $X$ contains noise, meaning I have observed $\hat{X} = X + u$ where $u$ represents the noise. When I run the OLS estimate, I find that the estimated $\beta$ is less than 1 due to the noise in $X$ , which inflates the variance of $X$ . I have different pairs of $(Y, X)$ from various groups. I can verify my assumption by observing that in larger groups (where data variability is reduced by the Central Limit Theorem), the fitted $\beta$ is closer to 1. Given this, my question is: If I want to make an out-of-sample forecast, should I use 1 as my $\beta$ or should I use the OLS $\beta$ estimate, which is biased? On one hand, it seems that using the unbiased $\beta$ of 1 is preferable since it is theoretically unbiased. However, if I assume the same noise $u$ is present in my out-of-sample $X$ , my residual term will have larger variance because it includes the term $\beta^2 \text{Var}(u)$ . Using a smaller, biased $\beta$ might help reduce the residual variance, despite the bias. Is this a bias-variance trade-off scenario? What is the best approach here if my goal is to minimize the out-of-sample residual MSE?
Quote
Report

Post Reply

Quoted from Forex.com.bd-Editorial External question — Cross Validated Stack Exchange Author: The One Source score (net votes, not local likes): 0 Original post: https://stats.stackexchange.com/questions/650835 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. I am working with a dataset that includes variables $Y$ and $X$ . I assume that $$ Y = \beta X + \epsilon $$ satisfies all the assumptions of OLS. Based on industry knowledge, I know that theoretically $\beta = 1$ . However, I am aware that $X$ contains noise, meaning I have observed $\hat{X} = X + u$ where $u$ represents the noise. When I run the OLS estimate, I find that the estimated $\beta$ is less than 1 due to the noise in $X$ , which inflates the variance of $X$ . I have different pairs of $(Y, X)$ from various groups. I can verify my assumption by observing that in larger groups (where data variability is reduced by the Central Limit Theorem), the fitted $\beta$ is closer to 1. Given this, my question is: If I want to make an out-of-sample forecast, should I use 1 as my $\beta$ or should I use the OLS $\beta$ estimate, which is biased? On one hand, it seems that using the unbiased $\beta$ of 1 is preferable since it is theoretically unbiased. However, if I assume the same noise $u$ is present in my out-of-sample $X$ , my residual term will have larger variance because it includes the term $\beta^2 \text{Var}(u)$ . Using a smaller, biased $\beta$ might help reduce the residual variance, despite the bias. Is this a bias-variance trade-off scenario? What is the best approach here if my goal is to minimize the out-of-sample residual MSE?

Cancel quote

Checking account access…