Is it better to use a MinMax or a Log Return normalization to predict stock price movements?
Is it better to use a MinMax or a Log Return normalization to predict stock price movements?
Loading saved threads...
Vincent Roye · External communityPost link
External question — Data Science Stack Exchange
Author: Vincent Roye
Original post: https://datascience.stackexchange.com/questions/40425
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
I am trying to use a LSTM model to predict d+2 and d+3 closing prices. I am not sure whether I should normalize the data
with a MixMax scaler (-1,+1)
using the log return
(P(n)-P(0))/P(0) for each sample
I have tried quite a lot of source code from Github and they don't seem to converge on any technique.
Quote
Report
keiv.fly · External communityPost link
External answer — Data Science Stack Exchange
Author: keiv.fly
Original post: https://datascience.stackexchange.com/a/40431
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
Log returns are symmetric compared to percentage change.
log(a/b) = - log(b/a)
and this (less skewness), in theory, leads to better results for most models (linear regression, neural networks).
Neural networks like lstm work better if the values are close to zero, but the difference in normalizations is usually not that big.
Any returns (log or percentage) are better than raw values because prices change according to previous prices. Their absolute (raw) values have almost negligible influence compared to previous price.
I would recommend first to convert to log returns and then normalize. If it is daily prices then I would divide the log returns by something like 0.05. Price changes have very heavy tails in a distribution so I would not suggest using minmax because then you divide by something like 0.5 (which probably was in great depression) and get all values too close to zero. Dividing by standard deviation should also be good.
But reality is different than theory, so it is better to benchmark. Maybe percentage changes are better because this is the number people see and react to. And markets are a lot about psychology.
And be prepared to see very high errors and bad models. Financial markets are badly predictable both in practice and theory. According to economic theory if they were predictable and people are rational and have unlimited credit lines then any possibility of earning additional money compared to the whole market will be closed in milliseconds. Only if you find some way to analyze data that noone is currently using only then will you be able to earn money. Neural networks were discussed in 1990s to predict financial markets. So LSTM is not really new in 2018.
Quote
Report
Georg M. Goerg · External communityPost link
External answer — Data Science Stack Exchange
Author: Georg M. Goerg
Original post: https://datascience.stackexchange.com/a/124266
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
I suggest to Gaussianize stock returns from heavy-tailed data to something that looks more like Gaussian. In particular, use Lambert W x Gaussian distributions to remove heavy-tails from data.
This is available in both in the
LambertW
R package and in Python's
pylambertw
module.
As an example of the ability to Gaussianize data with very heavy-tails consider the example of removing heavy-tails in a standard Cauchy random sample (a standard Cauchy does not even have a well-defined mean).
import pylambertw
from pylambertw.utils import plot
import numpy as np
np.random.seed(42)
y = np.random.standard_cauchy(size=1000)
plot.test_norm(y)
Then you can use a Lambert W x Gaussian methods of moments (IGMM) estimator to train a transformer that normalizes the data.
import pylambertw.igmm
clf = pylambertw.igmm.IGMM()
clf.fit(y)
x = clf.transform(y)
plot.test_norm(x)
See also the
Gaussianizer()
transformer to operate on multi-dimensional X.
See Goerg (2011 & 2015) for the original papers, with a detailed application of this methodology on stock return data [removing skewness & heavy-tails].
Quote
Report
Post Reply
Quoted from Forex.com.bd-Editorial External answer — Data Science Stack Exchange Author: Georg M. Goerg Source score (net votes, not local likes): 0 Original post: https://datascience.stackexchange.com/a/124266 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. I suggest to Gaussianize stock returns from heavy-tailed data to something that looks more like Gaussian. In particular, use Lambert W x Gaussian distributions to remove heavy-tails from data. This is available in both in the LambertW R package and in Python's pylambertw module. As an example of the ability to Gaussianize data with very heavy-tails consider the example of removing heavy-tails in a standard Cauchy random sample (a standard Cauchy does not even have a well-defined mean). import pylambertw from pylambertw.utils import plot import numpy as np np.random.seed(42) y = np.random.standard_cauchy(size=1000) plot.test_norm(y) Then you can use a Lambert W x Gaussian methods of moments (IGMM) estimator to train a transformer that normalizes the data. import pylambertw.igmm clf = pylambertw.igmm.IGMM() clf.fit(y) x = clf.transform(y) plot.test_norm(x) See also the Gaussianizer() transformer to operate on multi-dimensional X. See Goerg (2011 & 2015) for the original papers, with a detailed application of this methodology on stock return data [removing skewness & heavy-tails].
Checking account access…