How do I deal with unbalance classes in a stock market prediction problem?

How do I deal with unbalance classes in a stock market prediction problem?

Manage alerts

Loading saved threads...

user3118602 · External communityPost link
External question — Data Science Stack Exchange Author: user3118602 Original post: https://datascience.stackexchange.com/questions/102496 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. I am working on a prediction model to predict whether a stock should sell, hold or buy in n days. Each day (or row in the dataset), I classify whether this should be sell, hold or buy based on the percentage change and a new column will be created to indicate what is the action for that particular day. How should I deal with unbalance classification in my dataset when training my model? The train set as it is looks like this: 1 1401 0 835 -1 413 # 1 is buy, 0 is hold, -1 is sell From reading up, balancing depends on the problem. Do I need to balance my data for a stock market prediction classification? Thanks in advance. PS: I am using SVM and Naive Bayes.
Quote
Report
serali · External communityPost link
External answer — Data Science Stack Exchange Author: serali Original post: https://datascience.stackexchange.com/a/102497 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. What percentages are you using for buy, hold and sell classes? From the data you share in the question, I am guessing it is a stock that has been going up rather than down for the most of the days. So, if you increase percentage cutoffs you have for the stock, you will have a balanced data. As you don't share the details in your question, let's assume you set your classes to signal "buy" if change is larger than %1, sell for lower than -%1 and hold anywhere in between. But if you set the "buy" cutoff to be -let's say- %2 and "sell" cutoff to be %0, you might end up with a better balanced data. To get the exact points which will give you balanced data, you can use quantiles method with q= 1/3 .
Quote
Report
Mateusz · External communityPost link
External answer — Data Science Stack Exchange Author: Mateusz Original post: https://datascience.stackexchange.com/a/102499 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. The usual approach with unbalanced classes is just to make the train and test sets as homogenous as possible. So make sure that proportions of the classes in both sets are the same. There are many factors that can be taken into account when splitting data, but I'm gonna guess that you just need the basic approach. In sklearn that would be any stratified sampling. To investigate if the class imbalances do not cause a problem you can then see if the model predicts some classes in the test set worse than the others. You could then adjust the thresholds for classifying samples to get rid of some imbalances. Although I would not suspect that to be the case with the two models you are using, but I might be wrong and it doesn't hurt to check. Also in Naive Bayes the proportions of classes are an informative input for the model. They are known as the priors. I think most libraries take care of calculating the priors themselves and you shouldn't change them, unless you have a good reason to do so.
Quote
Report
Dave · External communityPost link
External answer — Data Science Stack Exchange Author: Dave Original post: https://datascience.stackexchange.com/a/112276 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. There is a perfect mapping between the rule and the decision. You know the percentage change; now apply your rule to map that to a buy/sell/hold decision. There is no machine learning to do, but even if there were, now is a nice time to post a reminder that class imbalance isn't much of a problem when proper statistical methods are used . Models like the SVM and naïve Bayes approaches are useful because they discover rules that you did not know. Logistic regression figures out the optimal coefficients, even if we specify the functional form. However, you already know the rule for mapping percent change to a buy/hold/sell decision. You don't have to figure it out.
Quote
Report

Post Reply

Quoted from Forex.com.bd-Editorial External question — Data Science Stack Exchange Author: user3118602 Source score (net votes, not local likes): 1 Original post: https://datascience.stackexchange.com/questions/102496 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. I am working on a prediction model to predict whether a stock should sell, hold or buy in n days. Each day (or row in the dataset), I classify whether this should be sell, hold or buy based on the percentage change and a new column will be created to indicate what is the action for that particular day. How should I deal with unbalance classification in my dataset when training my model? The train set as it is looks like this: 1 1401 0 835 -1 413 # 1 is buy, 0 is hold, -1 is sell From reading up, balancing depends on the problem. Do I need to balance my data for a stock market prediction classification? Thanks in advance. PS: I am using SVM and Naive Bayes.

Cancel quote

Checking account access…