Should I deal with missing values first then transform the data or vice versa?
Should I deal with missing values first then transform the data or vice versa?
Loading saved threads...
MINH NHỰT NGUYỄN TRẦN · External communityPost link
External question — Data Science Stack Exchange
Author: MINH NHỰT NGUYỄN TRẦN
Original post: https://datascience.stackexchange.com/questions/113922
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
I am currently working on a project involving time series banking stock price data. I have around 3000 observations, some columns have a lot of missing values (null value); they can account for 5 to 50% of the total observations. I have no idea what is the proper order for handling missing values, outliers and take log transformation of the data. Should I impute the missing values first and take log transformation or vice versa. Thank you so much
Quote
Report
jeffhale · External communityPost link
External answer — Data Science Stack Exchange
Author: jeffhale
Original post: https://datascience.stackexchange.com/a/113926
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
I suggest imputing the missing values and converting any columns to numeric data types as your first steps. Then you can deal with outliers and make transformations.
Most machine learning packages (e.g. scikit-learn) will generally require you to have all numeric data before feeding the data to machine learning algorithms.
As you go, just make sure you document what you're doing so that other people can follow and share you code so you can reproduce your work.
Quote
Report
Nicolas Martin · External communityPost link
External answer — Data Science Stack Exchange
Author: Nicolas Martin
Original post: https://datascience.stackexchange.com/a/113932
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
In general, it is better to deal with missing values first because there could be data loss or additional noise applying operations like a log that could impact classification or prediction algorithms.
To deal with missing values, you can use regressors to have good results but it depends on the data quality.
It could be done using algorithms such as Random Forest, XGBoost or Deep Neural Networks.
Note: you can measure the model quality by hiding some known values and see if they are well predicted.
See also:
https://cardoai.com/handling-missing-data-with-python/
Quote
Report
Ching · External communityPost link
External answer — Data Science Stack Exchange
Author: Ching
Original post: https://datascience.stackexchange.com/a/129737
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
From this book, it actually claims that transform and then impute could potentially be better as it preserves the variables relationship, e.g. Box-Cox transformation could be distorted if transformation is after.
For stock data, its log-normal, and when thinking about ffilling the data first, and log-diff to get returns, seems harmless as its basically saying on weekends return is 0.
For reference,
https://www.bookdown.org/rwnahhas/RMPH/mi-fitting.html
Quote
Report
Post Reply
Quoted from Forex.com.bd-Editorial External answer — Data Science Stack Exchange Author: jeffhale Source score (net votes, not local likes): 0 Original post: https://datascience.stackexchange.com/a/113926 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. I suggest imputing the missing values and converting any columns to numeric data types as your first steps. Then you can deal with outliers and make transformations. Most machine learning packages (e.g. scikit-learn) will generally require you to have all numeric data before feeding the data to machine learning algorithms. As you go, just make sure you document what you're doing so that other people can follow and share you code so you can reproduce your work.
Checking account access…