Can a biased hypothesis test be preferred over an unbiased one?

Can a biased hypothesis test be preferred over an unbiased one?

Manage alerts

Loading saved threads...

Shirin Ahmadov · External communityPost link
External question — Cross Validated Stack Exchange Author: Shirin Ahmadov Original post: https://stats.stackexchange.com/questions/664822 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. In point estimation, a biased estimator can be preferred over an unbiased one considering bias-variance trade-off. Is this also the case for hypothesis testing? More generally, can a biased hypothesis be accepted?
Quote
Report
Preston Botter · External communityPost link
External answer — Cross Validated Stack Exchange Author: Preston Botter Original post: https://stats.stackexchange.com/a/664827 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. In point estimation, there's a well-known bias-variance trade-off. Sometimes, a biased estimator may be preferred over an unbiased one if it substantially reduces variance $Var(\hat{\theta})$ , thus lowering mean squared error (MSE): $MSE(\hat{\theta})$ + $Var(\hat{\theta})$ + $Bias[(\hat{\theta})]^2$ , where $\theta$ is the true but unknown parameter, and is a statistic computed from a sample, used to estimate $\hat{\theta}$ . So, yes—biased estimators can be desirable. Bayesian Example A good example of this comes from Bayesian estimation using an informative prior. Suppose prior research suggests that a treatment effect $\theta$ is likely around 2 units, and we encode this belief with a prior distribution $\theta \sim N(2,1^2)$ . Given new sample data, we update this prior using Bayes’ rule, yielding a posterior distribution that combines both the data and prior information. The resulting Bayesian estimator (e.g., the posterior mean) $\hat{\theta}_{Bayes}=w*\hat{\theta}_{MLE}+(1-w)*2$ shrinks the estimate toward the prior mean, with the weight $w$ depending on the sample size and variance. This introduces bias (since the estimator is no longer centered solely on the data), but can substantially reduce variance—especially with small samples—resulting in a lower overall MSE. Thus, the informative prior provides a principled way to accept bias in exchange for greater estimation accuracy.
Quote
Report
Peter Flom · External communityPost link
External answer — Cross Validated Stack Exchange Author: Peter Flom Original post: https://stats.stackexchange.com/a/664828 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. There are several cases where a biased parameter estimate may be preferred to an unbiased one. Preston gives a Bayesian example. One non-Bayesian example is ridge regression, which is often recommended when there is collinearity. The ridge parameter estimate is given by $\beta_r = \text{argmin}_{\beta}(y-X\beta)^t(y-X\beta) + \lambda(\beta^t \beta -c)$ Where c is a constant. A more prosaic example is a bathroom scale, where a scale that is (say) biased downwards by 1 kg but has an sd of 0.1 kg may be preferred to one that is unbiased, but has sd of 1.0 kg. I should think these biased estimators lead to biased hypothesis tests.
Quote
Report
Christian Hennig · External communityPost link
External answer — Cross Validated Stack Exchange Author: Christian Hennig Original post: https://stats.stackexchange.com/a/664829 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. Biasedness in hypothesis testing means that there are distributions in the alternative for which the rejection probability is smaller than or equal to the significance level $\alpha$ . Let's say we run a parametric test regarding a parameter $\theta$ with $H_0:\ \theta=\theta_0,\ H_1:\ \theta\neq\theta_0$ . Let $R$ be the event of rejection and $A=R^c$ be the event of non-rejection. Now assume $P_{\theta_0}(R)=\alpha,\ \exists\theta^*:\ P_{\theta^*}(R)\le\alpha$ . Regarding your second question, like many statisticians I don't like the "accept" wording anyway. Obviously not rejecting $H_0$ doesn't mean it's true, be the test biased or not. If we observe $A$ , it's an event with a high probability $1-\alpha$ under $H_0$ , meaning that the data don't indicate against $H_0$ , $H_0$ is not rejected, and that makes as much sense as if the test were unbiased, because biasedness just doesn't change this. Now some people say "accept" in such a case, which I don't like, but that's not in the first place because of the bias of the test. What can be said is that the test doesn't distinguish $\theta_0$ from $\theta^*$ , as the data are just as much in line with $\theta^*$ as with $\theta_0$ . Of course this may be a reason for somebody to not say "accept" who'd otherwise say it, which is fair enough, but I'd then rather try to convince this person to not use this term in any case. To the other question: Can a biased test be preferable? Test quality is measured by power, and the power depends on $\theta$ . By definition of bias, the power detecting $\theta^*$ is bad, so that is a disadvantage. But it can absolutely be that a biased test has good power against other $\theta\neq\theta_0,\ \theta\neq\theta^*$ . Here is a somewhat stupid example. The two-sided t-test for a Gaussian mean $\mu$ is unbiased against the two-sided alternative $\mu\neq\mu_0$ with $H_0:\ \mu=\mu_0$ . The one-sided t-test at the same level (say the one that only rejects for larger $t$ indicating $\mu>\mu_0$ ) is a valid but biased test against the alternative $ \mu\neq\mu_0$ . If $\mu>\mu_0$ it has better power, in case $\mu<\mu_0$ it has power worse that $\alpha$ . I think it was Neyman who called such a test "worse than useless" (as non-rejection is even more likely for some members of $H_1$ than under $H_0$ ), but for some $\mu$ it is useful as the power is better than the power of the two-sided test. Now you might say, "but the one-sided test isn't biased - it is unbiased against its alternative, which is $\mu>\mu_0$ , not $\mu\neq\mu_0$ ". As it is defined, however, it is a valid test for $H_0:\ \mu=\mu_0$ regardless of the alternative (note that Fisher would define tests without reference to an alternative, and some follow him in this respect). Whether it is unbiased or not depends the formulation of the alternative. Personally I'd use the terminology like this: I'd call all distributions $Q$ with $Q(R)>\alpha$ the effective alternative of a test. These are all those distributions that the test distinguishes from the null hypothesis, in the sense of having a larger probability to observe $R$ . With this terminology, every test is unbiased against its effective alternative and biased tests no longer are a thing. Whoever uses a biased test needs to have in mind that some things in the nominal alternative are not in the effective alternative, and that the test basically has no power to detect these. If you intend to have power also against these, choose a different test. However, in some cases you may be happy with larger power against some members of the alternative and may be up for sacrificing other members. That's basically the relevant consideration. Chances are there are biased tests in some situations that are a little bit biased against a small bit of the alternative, and it is for mathematical or computational reasons impossible to have an unbiased test that is any good. In such cases people would use the biased test as they don't have a better one, but still it'd be good to know that and how it is biased, i.e., which parameter values are not really part of the effective alternative . I don't have an example in my mind right now but I believe such situations exist. Chances a good number of asymptotically valid tests are a little bit biased for finite samples.
Quote
Report

Post Reply

Quoted from Forex.com.bd-Editorial External answer — Cross Validated Stack Exchange Author: Christian Hennig Source score (net votes, not local likes): 5 Original post: https://stats.stackexchange.com/a/664829 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. Biasedness in hypothesis testing means that there are distributions in the alternative for which the rejection probability is smaller than or equal to the significance level $\alpha$ . Let's say we run a parametric test regarding a parameter $\theta$ with $H_0:\ \theta=\theta_0,\ H_1:\ \theta\neq\theta_0$ . Let $R$ be the event of rejection and $A=R^c$ be the event of non-rejection. Now assume $P_{\theta_0}(R)=\alpha,\ \exists\theta^*:\ P_{\theta^*}(R)\le\alpha$ . Regarding your second question, like many statisticians I don't like the "accept" wording anyway. Obviously not rejecting $H_0$ doesn't mean it's true, be the test biased or not. If we observe $A$ , it's an event with a high probability $1-\alpha$ under $H_0$ , meaning that the data don't indicate against $H_0$ , $H_0$ is not rejected, and that makes as much sense as if the test were unbiased, because biasedness just doesn't change this. Now some people say "accept" in such a case, which I don't like, but that's not in the first place because of the bias of the test. What can be said is that the test doesn't distinguish $\theta_0$ from $\theta^*$ , as the data are just as much in line with $\theta^*$ as with $\theta_0$ . Of course this may be a reason for somebody to not say "accept" who'd otherwise say it, which is fair enough, but I'd then rather try to convince this person to not use this term in any case. To the other question: Can a biased test be preferable? Test quality is measured by power, and the power depends on $\theta$ . By definition of bias, the power detecting $\theta^*$ is bad, so that is a disadvantage. But it can absolutely be that a biased test has good power against other $\theta\neq\theta_0,\ \theta\neq\theta^*$ . Here is a somewhat stupid example. The two-sided t-test for a Gaussian mean $\mu$ is unbiased against the two-sided alternative $\mu\neq\mu_0$ with $H_0:\ \mu=\mu_0$ . The one-sided t-test at the same level (say the one that only rejects for larger $t$ indicating $\mu>\mu_0$ ) is a valid but biased test against the alternative $ \mu\neq\mu_0$ . If $\mu>\mu_0$ it has better power, in case $\mu<\mu_0$ it has power worse that $\alpha$ . I think it was Neyman who called such a test "worse than useless" (as non-rejection is even more likely for some members of $H_1$ than under $H_0$ ), but for some $\mu$ it is useful as the power is better than the power of the two-sided test. Now you might say, "but the one-sided test isn't biased - it is unbiased against its alternative, which is $\mu>\mu_0$ , not $\mu\neq\mu_0$ ". As it is defined, however, it is a valid test for $H_0:\ \mu=\mu_0$ regardless of the alternative (note that Fisher would define tests without reference to an alternative, and some follow him in this respect). Whether it is unbiased or not depends the formulation of the alternative. Personally I'd use the terminology like this: I'd call all distributions $Q$ with $Q(R)>\alpha$ the effective alternative of a test. These are all those distributions that the test distinguishes from the null hypothesis, in the sense of having a larger probability to observe $R$ . With this terminology, every test is unbiased against its effective alternative and biased tests no longer are a thing. Whoever uses a biased test needs to have in mind that some things in the nominal alternative are not in the effective alternative, and that the test basically has no power to detect these. If you intend to have power also against these, choose a different test. However, in some cases you may be happy with larger power against some members of the alternative and may be up for sacrificing other members. That's basically the relevant consideration. Chances are there are biased tests in some situations that are a little bit biased against a small bit of the alternative, and it is for mathematical or computational reasons impossible to have an unbiased test that is any good. In such cases people would use the biased test as they don't have a better one, but still it'd be good to know that and how it is biased, i.e., which parameter values are not really part of the effective alternative . I don't have an example in my mind right now but I believe such situations exist. Chances a good number of asymptotically valid tests are a little bit biased for finite samples.

Cancel quote

Checking account access…