Incorrect Text Classification, But Accurate Model. Do I Perform Manual Text Classification For A Data Set?
Incorrect Text Classification, But Accurate Model. Do I Perform Manual Text Classification For A Data Set?
Loading saved threads...
BorangeOrange1337 · External communityPost link
External question — Data Science Stack Exchange
Author: BorangeOrange1337
Original post: https://datascience.stackexchange.com/questions/53462
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
I'm currently using Google's BERT pre-trained sentiment analysis model that is trained on an IMDb pos/neg review dataset. I'm using this model to predict whether tweets are positive (bullish) or negative (bearish). While the model is accurate when plugging in my own test data (F1 Score ~86%), the classification itself is not accurate. Tweets that are undoubtedly positive/bullish, and not classified as so. Perhaps this is because the language in the investment world is different than a movie review - which uses universally recognized positive/negative words and/or sentences.
The same is true when I take my tweet dataset and use Vader SentimentIntensityAnalyser to parse pos/neg tweets into separate folders.
So my question is... since the language that is used for telling whether a stock is bullish/bearish is uniquely different from that of an Amazon review, or movie review, would it be optimal for me to manually classify my dataset into positive (bullish) and negative (bearish) datasets?
Quote
Report
Erwan · External communityPost link
External answer — Data Science Stack Exchange
Author: Erwan
Original post: https://datascience.stackexchange.com/a/53469
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
There can be two distinct reasons to use instances annotated with the gold-standard class, i.e. the true answer for the target application:
In order to perform proper
evaluation
your test set must contain the gold-standard labels. The principle of evaluation is to measure by how much the predictions deviate from the truth, but without the truth the performance that you obtain on the test set is meaningless for the task that you are doing.
In order to train a
supervised
or semi-supervised model, the training set must contain the gold-standard labels. Semi-supervised methods offer some options to adapt a training set to a different task.
You can't rely on a model if you can't evaluate it at least on a small sample, so yes you probably need to manually annotate a subset of the data. It's only after that you can start thinking about how to improve performance.
Quote
Report
Post Reply
Quoted from Forex.com.bd-Editorial External answer — Data Science Stack Exchange Author: Erwan Source score (net votes, not local likes): 1 Original post: https://datascience.stackexchange.com/a/53469 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. There can be two distinct reasons to use instances annotated with the gold-standard class, i.e. the true answer for the target application: In order to perform proper evaluation your test set must contain the gold-standard labels. The principle of evaluation is to measure by how much the predictions deviate from the truth, but without the truth the performance that you obtain on the test set is meaningless for the task that you are doing. In order to train a supervised or semi-supervised model, the training set must contain the gold-standard labels. Semi-supervised methods offer some options to adapt a training set to a different task. You can't rely on a model if you can't evaluate it at least on a small sample, so yes you probably need to manually annotate a subset of the data. It's only after that you can start thinking about how to improve performance.
Checking account access…