Does bias in statistics and machine learning mean the same thing?

Does bias in statistics and machine learning mean the same thing?

Manage alerts

Loading saved threads...

kokoma · External communityPost link
External question — Cross Validated Stack Exchange Author: kokoma Original post: https://stats.stackexchange.com/questions/256447 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. In statistics, people often talk about unbiased estimators. In machine learning, bias variance trade-off is mentioned all the time. Does bias in both contexts mean the same thing? Does an unbiased estimator have a bias for the data it tries to model?
Quote
Report
yz616 · External communityPost link
External answer — Cross Validated Stack Exchange Author: yz616 Original post: https://stats.stackexchange.com/a/256592 License: CC BY-SA 3.0 — https://creativecommons.org/licenses/by-sa/3.0/ Adaptation: HTML converted to plain text; contact email addresses removed. Yes, they mean the same thing. This free chapter covers bias and variance of estimators: http://www.deeplearningbook.org/contents/ml.html Please see section 5.4, which has a good explanation of what they are.
Quote
Report
amorim-ds · External communityPost link
External answer — Cross Validated Stack Exchange Author: amorim-ds Original post: https://stats.stackexchange.com/a/455733 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. No, they don't. But they're similar. In ML the learning bias is the set of wrong assumptions that a model makes to fit a dataset. That can be thought of as a measure of how well the model fits the training dataset. On the other hand, the regular statistic bias is mathematically defined as the average of the absolute error of an estimator. If this number is zero the estimator (or model) is unbidden, if it is positive then the estimator is positive biased, which means the on average the estimation (or predictions) will be always higher than the true value. They're different because a model can have learning bias (it doesn't fit perfectly the data set) but it is unbiased (the average absolute error is equal to zero, or very close to zero) . For a more detailed explanation read this paper: http://www.cems.uwe.ac.uk/~irjohnso/coursenotes/uqc832/tr-bias.pdf
Quote
Report
Victor Kostyuk · External communityPost link
External answer — Cross Validated Stack Exchange Author: Victor Kostyuk Original post: https://stats.stackexchange.com/a/648769 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. The concepts are related but not the same, as in ML bias refers to a specific instance of more general statistical bias. Both terms refer to an error -- i.e., the difference between the expected value and the real value. However, in the ML setting, it refers specifically to error in estimating the value in the training data. In statistics context, it usually refers to error in an estimator of some statistic as the sample size (randomly sampled) increases to infinity. Bias in statistics refers to whether an expected value of an estimator (e.g., as number of samples goes to infinity) is equal to the quantity being estimated. E.g., the sample mean is an unbiased estimator of the population mean if the samples are random. The bias is the difference between the expected value of the estimator and the quantity estimated. In ML, the bias in bias-variance tradeoff refers specifically to the difference between expected value of the model and the actual label value for points in the training data (not all randomly sampled data as in general for statistical bias). Thus, if the model, after training, is able to perfectly predict all training data, it has 0 bias (and considerable variance, if the training data labels have variance -- hence the "tradeoff"). That's usually not good, because it typically means that the model is overfitted to the training data. Vice versa, if the model has significant errors when applied to the training data, it is high bias. For example, linear regression would have high bias when applied to training data generated by a non-linear (e.g., polynomial) process. This is where the divergence in terminology can be confusing: regression (OLS) is an unbiased estimator of coefficients (large samples will have increasingly similar fit to running regression on the population as a whole), but will likely be biased in estimating the labels in a selected training data sample. A model is underfit if it poorly predicts the training data. That is, it has high bias in the bias-variance tradeoff. It can be due to insufficiently rich model structure (as in the example above of a linear model predicting a polynomial pattern), or it can be due to insufficient training (e.g., only 1 epoch in training your NN), or non-predictive features (i.e., no relationship between features and label).
Quote
Report

Post Reply

Quoted from Forex.com.bd-Editorial External answer — Cross Validated Stack Exchange Author: amorim-ds Source score (net votes, not local likes): -2 Original post: https://stats.stackexchange.com/a/455733 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. No, they don't. But they're similar. In ML the learning bias is the set of wrong assumptions that a model makes to fit a dataset. That can be thought of as a measure of how well the model fits the training dataset. On the other hand, the regular statistic bias is mathematically defined as the average of the absolute error of an estimator. If this number is zero the estimator (or model) is unbidden, if it is positive then the estimator is positive biased, which means the on average the estimation (or predictions) will be always higher than the true value. They're different because a model can have learning bias (it doesn't fit perfectly the data set) but it is unbiased (the average absolute error is equal to zero, or very close to zero) . For a more detailed explanation read this paper: http://www.cems.uwe.ac.uk/~irjohnso/coursenotes/uqc832/tr-bias.pdf

Cancel quote

Checking account access…