Non IID variables and SVM Classifier
Non IID variables and SVM Classifier
Loading saved threads...
Aditya Kulkarni · External communityPost link
External question — Data Science Stack Exchange
Author: Aditya Kulkarni
Original post: https://datascience.stackexchange.com/questions/94344
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
I am training an SVM model to predict the trend of stock prices (one-day ahead predictions. Classification task). It Had completely slipped from my mind that SVMs assume IID data until I had a conversation with a friend.
This made me rethink about my approach and I have a few questions.
Why does SVM assume IID data in the first place?
By IID, does it mean that all the features should be linearly uncorrelated or should there not be any non-linear relationships as well?
The features include basic stock data (Open, High, Low, Close, Volume) and technical indicators (derived from basic data). Of course, the technical indicators will have dependencies with basic data. So I removed the basic data from the feature space. However, this does mean that the technical indicators are not related. Also,
How does one make sure that the data is identically distributed? Does scaling (Normalization or Standardization) help in this case?
Lastly, I have read some papers wherein people have used NN and SVM (both assume IID data) for stock trend prediction, but nowhere did I see any mention of the IID assumption and the fact that stock data is actually not uncorrelated and to some extent exhibits some autocorrelation as well. So,
How does one justify the use of non IID data as input to such algorithms which assume that the data is IID?
Thank you :)
Quote
Report
Abhishek Verma · External communityPost link
External answer — Data Science Stack Exchange
Author: Abhishek Verma
Original post: https://datascience.stackexchange.com/a/94351
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
The assumption regarding IID variables is to ensure a unique solution. By IID, it means they should be uncorrelated. You can't make sure data is identically distributed. Scaling and standardization are the obvious for you to find the solution. One justifies the use by getting satisfied by the solution found by the SVM algorithm (which is just one of the solution it has found).
But, you can the parameters for SVM and use a different kernel altogether. Remove redundant features yourself.
Or if you are looking at stock price prediction, look at Jane Market Street Prediction on Kaggle. I have seen heavy use of trees, mixture models and NNs.
Quote
Report
Post Reply
Quoted from Forex.com.bd-Editorial External answer — Data Science Stack Exchange Author: Abhishek Verma Source score (net votes, not local likes): 0 Original post: https://datascience.stackexchange.com/a/94351 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. The assumption regarding IID variables is to ensure a unique solution. By IID, it means they should be uncorrelated. You can't make sure data is identically distributed. Scaling and standardization are the obvious for you to find the solution. One justifies the use by getting satisfied by the solution found by the SVM algorithm (which is just one of the solution it has found). But, you can the parameters for SVM and use a different kernel altogether. Remove redundant features yourself. Or if you are looking at stock price prediction, look at Jane Market Street Prediction on Kaggle. I have seen heavy use of trees, mixture models and NNs.
Checking account access…