Difference between transforming individual features and taking their polynomial transformations?
Difference between transforming individual features and taking their polynomial transformations?
Loading saved threads...
plotmaster473 · External communityPost link
External question — Cross Validated Stack Exchange
Author: plotmaster473
Original post: https://stats.stackexchange.com/questions/670647
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
I am seeking clarity around the three concepts of:
Feature engineering on non-normal variables;
Creating polynomial transformations of variables where the relationship appears to be non-linear;
[Edited: this is in the context of] the computational cost of such transformations in the context of a broader machine learning pipeline.
As an example, I have a small dummy dataset of shape (2000, 14) on which I am performing multinomial (softmax) logistic regression. This dataset has some non-normality present in certain variables, as well as some polynomial relationships between variables. See figs 1 and 2 below.
[fig. 1]
[fig. 2]
Creating Polynomial transformations by iterating through several degrees creates coefficients and variables for all interactions, which is exhaustive in nature, but I am not sure this is what I want, as not all variables have polynomial relationships. This is somewhat of a moot point, as the coefficients for interaction terms that are not relevant will tend towards zero -- but I am curious (firstly) if there is any statistical preference for doing this exhaustively (as there may be some hidden patterns that are invisible to a quick-peek EDA) compared with using ColumnTransformer on a subset where it is clear that polynomial interactions exist.
Secondly, I am completely ignorant of how transforming individual features from non-normal to normal plays into this. I am familiar with feature engineering, but I wonder if the community can help me better understand how transforming individual features to normality would affect the polynomial transformations needed for the model to optimize the coefficients.
Thirdly, I worry about the computational cost. If I perform exhaustive search feature selection while looping through polynomials, compute time explodes -- and this is on a dummy dataset. I feel there must be a way to control for the polynomial transformations.
To summarize:
Will transforming ALL the features somehow obfuscates the features that are normally distributed / have linear relationships with one another? What is the cost-benefit trade-off with this?
Should I be transforming individual variables to have normal distributions (read: creating normal, non-skewed distributions; this does NOT mean standardizing) in addition to polynomial transformations?
Is there any "trick" to getting around the computational cost of performing feature selection in tandem with polynomial transformations? Or is optimizing in other ways preferable?
Thanks in advance.
Quote
Report
EdM · External communityPost link
External answer — Cross Validated Stack Exchange
Author: EdM
Original post: https://stats.stackexchange.com/a/670674
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
Briefly:
Predictor variables do not need to be normally distributed, even in simple linear regression. See
this page
. That should help with your Question 2.
Trying to fit a single polynomial across the full range of a predictor will tend to lead to problems unless there is a solid theoretical basis for a particular polynomial form. A regression spline or some other type of generalized additive model is a much better choice. See
this answer
and others on that page. You can then check the statistical and practical significance of the nonlinear terms. That should help with Question 1.
Automated model selection is
not a good idea
. An exhaustive search for all possible interactions among (potentially transformed) predictors runs a big risk of overfitting. It's best to use your knowledge of the subject matter to include interactions that make sense. With a large data set, you could include a number of interactions that is unlikely to lead to overfitting based on your number of observations.
Chapter 4 of Regression Modeling Strategies
by Frank Harrell provides helpful guidance. Pre-specifying the model will help with Question 3, as you will no longer need to fit multiple models. If you really need additional "feature selection," consider limited backward elimination from the full model, making sure to keep all lower-level terms included in your retained interactions. See
this answer
.
Quote
Report
Post Reply
Quoted from Forex.com.bd-Editorial External question — Cross Validated Stack Exchange Author: plotmaster473 Source score (net votes, not local likes): 4 Original post: https://stats.stackexchange.com/questions/670647 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. I am seeking clarity around the three concepts of: Feature engineering on non-normal variables; Creating polynomial transformations of variables where the relationship appears to be non-linear; [Edited: this is in the context of] the computational cost of such transformations in the context of a broader machine learning pipeline. As an example, I have a small dummy dataset of shape (2000, 14) on which I am performing multinomial (softmax) logistic regression. This dataset has some non-normality present in certain variables, as well as some polynomial relationships between variables. See figs 1 and 2 below. [fig. 1] [fig. 2] Creating Polynomial transformations by iterating through several degrees creates coefficients and variables for all interactions, which is exhaustive in nature, but I am not sure this is what I want, as not all variables have polynomial relationships. This is somewhat of a moot point, as the coefficients for interaction terms that are not relevant will tend towards zero -- but I am curious (firstly) if there is any statistical preference for doing this exhaustively (as there may be some hidden patterns that are invisible to a quick-peek EDA) compared with using ColumnTransformer on a subset where it is clear that polynomial interactions exist. Secondly, I am completely ignorant of how transforming individual features from non-normal to normal plays into this. I am familiar with feature engineering, but I wonder if the community can help me better understand how transforming individual features to normality would affect the polynomial transformations needed for the model to optimize the coefficients. Thirdly, I worry about the computational cost. If I perform exhaustive search feature selection while looping through polynomials, compute time explodes -- and this is on a dummy dataset. I feel there must be a way to control for the polynomial transformations. To summarize: Will transforming ALL the features somehow obfuscates the features that are normally distributed / have linear relationships with one another? What is the cost-benefit trade-off with this? Should I be transforming individual variables to have normal distributions (read: creating normal, non-skewed distributions; this does NOT mean standardizing) in addition to polynomial transformations? Is there any "trick" to getting around the computational cost of performing feature selection in tandem with polynomial transformations? Or is optimizing in other ways preferable? Thanks in advance.
Checking account access…