Which preprocessing is the correct way to forecast time-series data using LSTM?
Which preprocessing is the correct way to forecast time-series data using LSTM?
Loading saved threads...
orde.r · External communityPost link
External question — Data Science Stack Exchange
Author: orde.r
Original post: https://datascience.stackexchange.com/questions/122018
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
I just started to study time-series forecasting using RNN.
I have a few months of time series data that was an hour unit.
The data is a kind of percentage value of my little experiment and no other correlated information this. It is simple 1-D array info.
I would like to forecast the future condition of this.
Many tutorials and web info introduced direct training and forecasting the time series data without any data pre-processing.
But for the RNN (or ML and DL), I think we should consider the data's condition that is stationary or not.
My data is totally random condition which is stationary data (no seasonality, no trend).
For example, the US stock prediction tutorial showed super great accuracy forecasting performance according to many LSTM tutorials.
[If this really works and is true, then all ML developers will be rich.]
And, Some of them didn't emphasize and note a kind of the data pre-processing such as non-stationary to stationary something like that.
According to my short knowledge, I think the non-stationary data such as stock price (will have trend) should be converted as a stationary format through differencing or some other steps. and I think this is a correct prediction as a view of theoretical sense even if the accuracy is not high.
So my point is, I'm a bit confused about whether that really is no need for any preprocessing to treat stationary or not.
For my case, I applied differencing step (
$t_n - t_{n-1}$
)to my time-series data in order to remove the trend or some periodic situation.
Is my understanding not correct?
Why do time-series forecasting tutorials not introduced data stationarity?
Quote
Report
Nicolas Martin · External communityPost link
External answer — Data Science Stack Exchange
Author: Nicolas Martin
Original post: https://datascience.stackexchange.com/a/122025
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
Models based on the stock markets are often unreliable because there is a lot of noise and even if it seems to predict well for the validation data, the result is very different in practice.
What you call stationary means depending on previous values in a relative way, and predicting stock markets should follow this rule.
However, RNN and LSTM have been built to memorize low-noise patterns, so you will want to apply some noise reduction like smoothing for better results.
In addition, I recommend converting data to have fluctuations (=derivates) rather than raw values, so that the neural network learns data dynamics.
For a good model, you must simulate prediction vs real-world results for each day. It means evaluating the model prediction for day 1, checking if the result is correct or not, then feeding the neural network with this information, and applying the same logic for day 2 and so on.
Generally speaking, stock markets are so difficult that you will want to use multi-variate and seasonality models. Maybe Prophet is a better option.
https://facebook.github.io/prophet/
Quote
Report
Post Reply
Quoted from Forex.com.bd-Editorial External answer — Data Science Stack Exchange Author: Nicolas Martin Source score (net votes, not local likes): 3 Original post: https://datascience.stackexchange.com/a/122025 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. Models based on the stock markets are often unreliable because there is a lot of noise and even if it seems to predict well for the validation data, the result is very different in practice. What you call stationary means depending on previous values in a relative way, and predicting stock markets should follow this rule. However, RNN and LSTM have been built to memorize low-noise patterns, so you will want to apply some noise reduction like smoothing for better results. In addition, I recommend converting data to have fluctuations (=derivates) rather than raw values, so that the neural network learns data dynamics. For a good model, you must simulate prediction vs real-world results for each day. It means evaluating the model prediction for day 1, checking if the result is correct or not, then feeding the neural network with this information, and applying the same logic for day 2 and so on. Generally speaking, stock markets are so difficult that you will want to use multi-variate and seasonality models. Maybe Prophet is a better option. https://facebook.github.io/prophet/
Checking account access…