How to scale a subset of data with respect to the entire dataset
How to scale a subset of data with respect to the entire dataset
Loading saved threads...
functorial · External communityPost link
External question — Data Science Stack Exchange
Author: functorial
Original post: https://datascience.stackexchange.com/questions/115363
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
I am developing a financial time-series prediction model using sklearn using
StandardScaler
for scaling purposes. I train a model, and then use the model regularly on data as it comes in. The training must be done in batches due to the large data size. Right now, I am scaling each batch using a different scaler for training each batch, and for each test/real data batch.
My concern is that removing the mean from a subset of the data removes a good deal of information - the scaled data for an asset trading at at \$1500 is not clearly distinguishable from the data for a penny stock.
I'm wondering whether it is possible to - and whether I should - continually re-fit the same scaler used in training, so that data is scaled to remove the mean
with respect to the entire dataset
(ie. train + the actual data being evaluated), and whether this is desirable.
Quote
Report
functorial · External communityPost link
External answer — Data Science Stack Exchange
Author: functorial
Original post: https://datascience.stackexchange.com/a/115364
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
I was able to find some consensus in various stack* sites - the answer is apparently to scale test data with respect to the train data - ie.
(testData - mean(trainData)) / sd(trainData)
, or with sklearn
scaler.fit(train_data); scaler.transform(test_data)
.
https://stats.stackexchange.com/questions/174823/how-to-apply-standardization-normalization-to-train-and-testset-if-prediction-i
Quote
Report
Post Reply
Quoted from Forex.com.bd-Editorial External answer — Data Science Stack Exchange Author: functorial Source score (net votes, not local likes): 1 Original post: https://datascience.stackexchange.com/a/115364 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. I was able to find some consensus in various stack* sites - the answer is apparently to scale test data with respect to the train data - ie. (testData - mean(trainData)) / sd(trainData) , or with sklearn scaler.fit(train_data); scaler.transform(test_data) . https://stats.stackexchange.com/questions/174823/how-to-apply-standardization-normalization-to-train-and-testset-if-prediction-i
Checking account access…