Why does precision decrease with inceasing threshold?
Why does precision decrease with inceasing threshold?
Loading saved threads...
Bryan Carty · External communityPost link
External question — Data Science Stack Exchange
Author: Bryan Carty
Original post: https://datascience.stackexchange.com/questions/128025
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
I've trained a Logistic Regression model using scikit-learns
LogisticRegression
class. I'm dealing with stock data so it's quite noisy and difficult to predict anything.
When graphing threshold vs. precision I can see that there's some degree of correlation but it's quite jagged and ultimately falls off as the threshold surpasses a certain point.
I'm wondering,
Why is this the case?
Any suggestions on how to increase prediction power - better suited model, an alternative metric to base prediction confidence off, better data preprocessing etc.
Any input would be greatly appreciated.
Quote
Report
RegressIt · External communityPost link
External answer — Data Science Stack Exchange
Author: RegressIt
Original post: https://datascience.stackexchange.com/a/128027
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
The jagged and falling-off nature of the precision-threshold graph in logistic regression can be attributed to several factors. One key factor is the nature of the data being used for the prediction. In the case of stock data, which is often noisy and volatile, the relationship between the input features and the target variable (stock movement, for example) may not be linear or easily discernible. This can lead to a jagged precision-threshold graph as the model struggles to make accurate predictions at different thresholds.
Another factor is the inherent limitations of logistic regression.
Logistic regression assumes a linear relationship between the input features and the log-odds of the target variable.
When this assumption is not met, especially in the presence of noisy and complex data such as stock movements, the model's predictive power can be limited, leading to a jagged precision-threshold graph.
Given the noisy and complex nature of stock data, it might be beneficial to explore alternative models that can capture non-linear relationships more effectively. Models such as
decision trees
,
random forests
, or even
neural networks
could potentially offer better predictive power in this context.
It's also important to conduct a thorough analysis of the input features and consider feature engineering techniques to extract more meaningful signals from the noisy stock data. Additionally, robust data preprocessing techniques, such as handling outliers and scaling features appropriately, can contribute to improved model performance.
You may want to also consider using alternative evaluation metrics such as F1 score, which balances precision and recall, especially in scenarios with imbalanced classes or noisy data.
If you still encounter issues you could explore the use of ensemble methods, such as
bagging
or
boosting
, to combine multiple models and potentially mitigate the impact of noise in the data.
Quote
Report
Post Reply
Quoted from Forex.com.bd-Editorial External question — Data Science Stack Exchange Author: Bryan Carty Source score (net votes, not local likes): 0 Original post: https://datascience.stackexchange.com/questions/128025 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. I've trained a Logistic Regression model using scikit-learns LogisticRegression class. I'm dealing with stock data so it's quite noisy and difficult to predict anything. When graphing threshold vs. precision I can see that there's some degree of correlation but it's quite jagged and ultimately falls off as the threshold surpasses a certain point. I'm wondering, Why is this the case? Any suggestions on how to increase prediction power - better suited model, an alternative metric to base prediction confidence off, better data preprocessing etc. Any input would be greatly appreciated.
Checking account access…