Should I deal with missing values first then transform the data or vice versa?

Should I deal with missing values first then transform the data or vice versa?

Manage alerts

Loading saved threads...

MINH NHỰT NGUYỄN TRẦN · External communityPost link
External question — Data Science Stack Exchange Author: MINH NHỰT NGUYỄN TRẦN Original post: https://datascience.stackexchange.com/questions/113922 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. I am currently working on a project involving time series banking stock price data. I have around 3000 observations, some columns have a lot of missing values (null value); they can account for 5 to 50% of the total observations. I have no idea what is the proper order for handling missing values, outliers and take log transformation of the data. Should I impute the missing values first and take log transformation or vice versa. Thank you so much
Quote
Report
jeffhale · External communityPost link
External answer — Data Science Stack Exchange Author: jeffhale Original post: https://datascience.stackexchange.com/a/113926 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. I suggest imputing the missing values and converting any columns to numeric data types as your first steps. Then you can deal with outliers and make transformations. Most machine learning packages (e.g. scikit-learn) will generally require you to have all numeric data before feeding the data to machine learning algorithms. As you go, just make sure you document what you're doing so that other people can follow and share you code so you can reproduce your work.
Quote
Report
Nicolas Martin · External communityPost link
External answer — Data Science Stack Exchange Author: Nicolas Martin Original post: https://datascience.stackexchange.com/a/113932 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. In general, it is better to deal with missing values first because there could be data loss or additional noise applying operations like a log that could impact classification or prediction algorithms. To deal with missing values, you can use regressors to have good results but it depends on the data quality. It could be done using algorithms such as Random Forest, XGBoost or Deep Neural Networks. Note: you can measure the model quality by hiding some known values and see if they are well predicted. See also: https://cardoai.com/handling-missing-data-with-python/
Quote
Report
Ching · External communityPost link
External answer — Data Science Stack Exchange Author: Ching Original post: https://datascience.stackexchange.com/a/129737 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. From this book, it actually claims that transform and then impute could potentially be better as it preserves the variables relationship, e.g. Box-Cox transformation could be distorted if transformation is after. For stock data, its log-normal, and when thinking about ffilling the data first, and log-diff to get returns, seems harmless as its basically saying on weekends return is 0. For reference, https://www.bookdown.org/rwnahhas/RMPH/mi-fitting.html
Quote
Report

Post Reply

Quoted from Forex.com.bd-Editorial External answer — Data Science Stack Exchange Author: Nicolas Martin Source score (net votes, not local likes): 4 Original post: https://datascience.stackexchange.com/a/113932 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. In general, it is better to deal with missing values first because there could be data loss or additional noise applying operations like a log that could impact classification or prediction algorithms. To deal with missing values, you can use regressors to have good results but it depends on the data quality. It could be done using algorithms such as Random Forest, XGBoost or Deep Neural Networks. Note: you can measure the model quality by hiding some known values and see if they are well predicted. See also: https://cardoai.com/handling-missing-data-with-python/

Cancel quote

Checking account access…