LSTM predicts the same value
LSTM predicts the same value
Loading saved threads...
user22615570 · External communityPost link
External question — Data Science Stack Exchange
Author: user22615570
Original post: https://datascience.stackexchange.com/questions/131427
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
I am implementing in PyTorch an LSTM model to predict if the closing value of a stock will go up or down in the next 5 and 10 minutes. Specifically, I am using 24 years of 5 minute data with 19 features, divided in chunks of one week per forecast (using 7 different stocks) The problem I'm facing is the fact that, no matter what, the LSTM model seems to predict values around one specific value to always minimize the loss, which it does not go down very much.
I pre-prepare the inputs and targets in torch.tensors with [batch_size, sequence_len, features] (which in my case is [32, 2016, 19]), normalize them between 0 and 1 and feed them to my LSTM model which is structured like this:
class MultiInputOutputLSTM(nn.Module):
def __init__(self, input_size, hidden_size, num_layers, output_size, dropout, lr, batch_size):
super(MultiInputOutputLSTM, self).__init__()
self.input_size = input_size
self.hidden_size = hidden_size
self.num_layers = num_layers
self.dropout = dropout
self.batch_size = batch_size
self.loss_list = []
self.accuracy = 0
self.predictions_list = [0]
self.lstm = nn.LSTM(input_size = self.input_size, hidden_size = self.hidden_size, num_layers = self.num_layers, dropout = self.dropout, batch_first=True)
self.fc = nn.Linear(hidden_size, output_size, bias=True)
self.sigmoid = nn.Sigmoid()
self.criterion = nn.BCEWithLogitsLoss()
self.optimizer = torch.optim.RMSprop(self.parameters(), lr = lr, alpha=0.9, weight_decay=1e-4, momentum=0.5)
self.scheduler = torch.optim.lr_scheduler.ReduceLROnPlateau(self.optimizer, 'min')
def forward(self, x):
h0 = torch.zeros(self.num_layers, self.batch_size, self.hidden_size)
c0 = torch.zeros(self.num_layers, self.batch_size, self.hidden_size)
lstm_out, _ = self.lstm(x, (h0, c0))
output = self.fc(lstm_out[:, -1, :])
return output
def train_step(self ,x, y):
self.train()
predictions_1 = torch.round(self.forward(x))
predictions = self.forward(x)
if (predictions_1.detach().cpu().numpy()[0] == y.detach().cpu().numpy()[0]).all():
self.accuracy += 1
self.predictions_list.append(predictions.detach().cpu().numpy()[0][0])
penalty = torch.mean((predictions-0.5)**2)
loss = self.criterion(predictions, y) + penalty
self.scheduler.step(loss)
self.optimizer.zero_grad()
self.loss_list.append(loss.item())
mean_loss = sum(self.loss_list)/len(self.loss_list)
loss.backward()
torch.nn.utils.clip_grad_norm_(self.parameters(), max_norm=6)
self.optimizer.step()
return loss.item(), mean_loss, self.accuracy, predictions.detach().cpu().numpy()[0], y.detach().cpu().numpy()[0]
The targets are 1 if the price goes up, 0 elsewise.
The hyper-parameters are:
input_size = 19
hidden_size = 3
num_layers = 5
output_size = 2
lr = 0.001
num_epochs = 5
batch_size = 32
dropout = 0.3
First, I'm creating the dataframe from the csv file: then I divide it in targets and inputs and with the latter I calculate the various trading signals.
I then transform the data in a suitable way for the LSTM model, dividing it in one-week chunks.
The model should predict values close to one if the price is going up and close to zero if the price stays the same/goes down: however it creates one prediction an then it will gradually go down to the 0.5 mark, staying still for the remainder of the learning process. The train-test split is 85-15%.
Here is a list of what I have tried: lowering or increasing learning rate (from 0.00001 to 0.1), output size (with more and less forecasts), batch_size (from 1 to 256), num_layers (1-5), input_size (with 1, 2, 3... 19 features), dropout (0 to 0.7) and num_epochs (from 1 to 100).
I have tried Adam optimiser, then SDG, then RMSProp optimiser changing alpha, momentum and weight_decay.
I have tried changing the data in input to better match the targets (substituted every element with 0 and 1: 1 if the element is bigger than the previous item, 0 else) or using the increments between elements instead.
Also tried BCELoss (with the sigmoid activation layer in self.forward()), L1Loss, MSELoss, CrossEntropyLoss, and even adjusting the loss adding a penalty
penalty = torch.mean((y-predictions)**(-2))
but it does not change anything, the values still float around 0.5. The mean loss goes down from 1.174 to 1.040.
I am now trying to penalise heavily values around 0.5 with
penalty = torch.mean((predictions-0.5)**(-2))
but the predictions go in one direction and stay around 0 or 1, not learning.
What can I do to resolve those issues? (someone suggested to me that it could be vanishing gradient, but I really don't know how to resolve this issue)
Quote
Report
Post Reply
Quoted from Forex.com.bd-Editorial External question — Data Science Stack Exchange Author: user22615570 Source score (net votes, not local likes): 1 Original post: https://datascience.stackexchange.com/questions/131427 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. I am implementing in PyTorch an LSTM model to predict if the closing value of a stock will go up or down in the next 5 and 10 minutes. Specifically, I am using 24 years of 5 minute data with 19 features, divided in chunks of one week per forecast (using 7 different stocks) The problem I'm facing is the fact that, no matter what, the LSTM model seems to predict values around one specific value to always minimize the loss, which it does not go down very much. I pre-prepare the inputs and targets in torch.tensors with [batch_size, sequence_len, features] (which in my case is [32, 2016, 19]), normalize them between 0 and 1 and feed them to my LSTM model which is structured like this: class MultiInputOutputLSTM(nn.Module): def __init__(self, input_size, hidden_size, num_layers, output_size, dropout, lr, batch_size): super(MultiInputOutputLSTM, self).__init__() self.input_size = input_size self.hidden_size = hidden_size self.num_layers = num_layers self.dropout = dropout self.batch_size = batch_size self.loss_list = [] self.accuracy = 0 self.predictions_list = [0] self.lstm = nn.LSTM(input_size = self.input_size, hidden_size = self.hidden_size, num_layers = self.num_layers, dropout = self.dropout, batch_first=True) self.fc = nn.Linear(hidden_size, output_size, bias=True) self.sigmoid = nn.Sigmoid() self.criterion = nn.BCEWithLogitsLoss() self.optimizer = torch.optim.RMSprop(self.parameters(), lr = lr, alpha=0.9, weight_decay=1e-4, momentum=0.5) self.scheduler = torch.optim.lr_scheduler.ReduceLROnPlateau(self.optimizer, 'min') def forward(self, x): h0 = torch.zeros(self.num_layers, self.batch_size, self.hidden_size) c0 = torch.zeros(self.num_layers, self.batch_size, self.hidden_size) lstm_out, _ = self.lstm(x, (h0, c0)) output = self.fc(lstm_out[:, -1, :]) return output def train_step(self ,x, y): self.train() predictions_1 = torch.round(self.forward(x)) predictions = self.forward(x) if (predictions_1.detach().cpu().numpy()[0] == y.detach().cpu().numpy()[0]).all(): self.accuracy += 1 self.predictions_list.append(predictions.detach().cpu().numpy()[0][0]) penalty = torch.mean((predictions-0.5)**2) loss = self.criterion(predictions, y) + penalty self.scheduler.step(loss) self.optimizer.zero_grad() self.loss_list.append(loss.item()) mean_loss = sum(self.loss_list)/len(self.loss_list) loss.backward() torch.nn.utils.clip_grad_norm_(self.parameters(), max_norm=6) self.optimizer.step() return loss.item(), mean_loss, self.accuracy, predictions.detach().cpu().numpy()[0], y.detach().cpu().numpy()[0] The targets are 1 if the price goes up, 0 elsewise. The hyper-parameters are: input_size = 19 hidden_size = 3 num_layers = 5 output_size = 2 lr = 0.001 num_epochs = 5 batch_size = 32 dropout = 0.3 First, I'm creating the dataframe from the csv file: then I divide it in targets and inputs and with the latter I calculate the various trading signals. I then transform the data in a suitable way for the LSTM model, dividing it in one-week chunks. The model should predict values close to one if the price is going up and close to zero if the price stays the same/goes down: however it creates one prediction an then it will gradually go down to the 0.5 mark, staying still for the remainder of the learning process. The train-test split is 85-15%. Here is a list of what I have tried: lowering or increasing learning rate (from 0.00001 to 0.1), output size (with more and less forecasts), batch_size (from 1 to 256), num_layers (1-5), input_size (with 1, 2, 3... 19 features), dropout (0 to 0.7) and num_epochs (from 1 to 100). I have tried Adam optimiser, then SDG, then RMSProp optimiser changing alpha, momentum and weight_decay. I have tried changing the data in input to better match the targets (substituted every element with 0 and 1: 1 if the element is bigger than the previous item, 0 else) or using the increments between elements instead. Also tried BCELoss (with the sigmoid activation layer in self.forward()), L1Loss, MSELoss, CrossEntropyLoss, and even adjusting the loss adding a penalty penalty = torch.mean((y-predictions)**(-2)) but it does not change anything, the values still float around 0.5. The mean loss goes down from 1.174 to 1.040. I am now trying to penalise heavily values around 0.5 with penalty = torch.mean((predictions-0.5)**(-2)) but the predictions go in one direction and stay around 0 or 1, not learning. What can I do to resolve those issues? (someone suggested to me that it could be vanishing gradient, but I really don't know how to resolve this issue)
Checking account access…