Improving the Accuracy of Sentiment Classification: A Novel Synthesis of Computational and Analytical Methods
CSEF · 2015 Mathematics & Software
Overview
Objectives/Goals The determination of individuals' mood in a review of a restaurant or the public sentiment of a political campaign are incredibly important statistics that cannot be accurately estimated manually or with simple computational techniques. The purpose of this project is to a) make use of supervised and weakly supervised machine learning algorithms to accurately classify the sentiment of a string of text, and b) incorporate negation handling, word n-grams, feature selection by mutual information, subjectivity classification, and polarity determination to improve classification of sentences. This project is unique as it is the first in the field to holistically explore a novel combination of both supervised and weakly-supervised machine learning models. Methods/Materials The two approaches studied were tested against corpora of data from IMDb, Amazon reviews, and Twitter for accuracy, precision, and recall. For each specified dataset, numerous iterations were run with different sample sizes ranging from 15 to 100. The primary analysis involved the use of IMDb pre-classified polar movie reviews. Every review was split into sentences, which were preprocessed, classified for subjectivity and polarity, and stored for future predictions. Results After training, feature selection by mutual information, and further textual analysis, the supervised model yielded an average accuracy of 88.7%. The weakly supervised model predictions continually increased in accuracy and were able to consistently predict results with an accuracy of greater than 83% after only 600 iterations. The weakly supervised model was more adept at making predictions on novel data due to its use of pattern matching and objectivity classification, whereas the supervised model prevailed at classifying sentences similar to its training set. Conclusions/Discussion The weakly supervised model improved on the foundations of the supervised model. The addition of subjectivity and polarity classification as well as feature selection vastly improved accuracies as only highly subjective sentences were included in overall calculations. The primary difference between the supervised and weakly supervised model was the analysis of linguistic patterns in sentences, which allowed for better classification of unseen cases.
Summary statement
This project compared and improved weakly supervised and supervised machine learning models using linguistic analysis, polarity and subjectivity classification, and negation handling to effectively classify the sentiment of provided text.
Help received
Parents helped with the board assembly. Computer Science teacher and mentor Dr. Eric Nelson helped with algorithm testing.
Competition history
- CSEF 2015
Resources
Related projects
ISEF · 2018
Automatically Analyzing Open-Ended Survey Responses Using Statistical and Machine Learning Methods
CSEF · 2017
Predicting Formality of Written Texts Using Machine Learning Algorithms
CSEF · 2014
A Machine Learning Model for Automated Semantic Short Essay Assessment through Random Forest Based Ensembles and NLP
CSEF · 2011
The Application of Bayesian Networks for Speech Classification
ISEF · 2015
#feels: Detecting and Visualizing Regional Sentiment from Cross-lingual Tweets for Specific Hashtags Using SVM and Naive Bayes Classifiers
CSEF · 2018
Analyzing Gender-Based Violence and Aggressive Behavior through Social Media Data
CSEF · 2019
Combating Cyberbullying and Toxicity by Teaching AI to Use Linguistic Insights from Human Interactions in Social Media
CSEF · 2017
Evaluation of Gender Bias in Social Media Using Artificial Intelligence
Closest projects by meaning, across every fair and year in the corpus.
Browse more like this
Source: California Science & Engineering Fair public projects