Analyzing Gender-Based Violence and Aggressive Behavior through Social Media Data
CSEF · 2018 Computational Systems & Analysis Second Award
Overview
Objectives/Goals My project goal was to develop a computational model to classify Tweets for relevance to gender-based violence (GBV), a prevalent human rights issue that transcends a variety of demographics. Methods/Materials I defined four classes that pertain to GBV: Physical Violence, Sexual Violence, Harmful Practices, and Other. I used the Python programming language. I extracted data from the Twitter API based on class-specific search criteria and employed the Natural Language Toolkit (NLTK) for natural language processing on Tweet text. Of the 4,000 filtered Tweets, 80% were used as training data and the rest were used as testing data. I used the Naive Bayes classification algorithm to train the machine learning model. I went on to conduct a comparative analysis of two feature sets consisting of unigrams and bigrams. I also constructed a confusion matrix to better analyze the model's performance. Results The feature set based on my search criteria had the highest accuracy, with over 85%. For the NLP-based features, Harmful Practices had the highest precision and Other had the lowest. For the search criteria-based features, Harmful Practices had the highest precision and Physical Violence had the lowest. The countries that most frequently discussed GBV included the US, UK, Canada, India, and Australia. Conclusions/Discussion I was able to meet my project goal, and successfully leveraged computational linguistics, machine learning, and computational social science to develop a highly accurate Tweet classification model for GBV.
Summary statement
I developed a computational model that can independently classify Tweets into one of four classes germane to gender-based violence (GBV) while harnessing natural language processing and machine learning.
Help received
I developed this model myself with the support of my parents and teacher sponsor.
Awards (1)
Competition history
- CSEF 2018
Resources
Related projects
CSEF · 2017
Predicting Formality of Written Texts Using Machine Learning Algorithms
CSEF · 2018
Cybian: A Machine Learning Program Based on Naive Bayes to Classify and Mitigate Cyberbullying
CSEF · 2015
Improving the Accuracy of Sentiment Classification: A Novel Synthesis of Computational and Analytical Methods
ISEF · 2015
Development of an Authorship Identification Algorithm for Twitter Using Stylometric Techniques
CSEF · 2017
Developing a Predictive Model for On-Campus Crime Using Machine Learning Algorithms and Reporting via Mobile App
CSEF · 2019
Combating Cyberbullying and Toxicity by Teaching AI to Use Linguistic Insights from Human Interactions in Social Media
ISEF · 2020
Using Natural Language Processing and Linguistic Insights to Combat Cyberbullying and Toxicity in Social Media
ISEF · 2017
Developing a Predictive Model for On-Campus Crime Using Machine Learning Algorithms and Reporting via Mobile App
Closest projects by meaning, across every fair and year in the corpus.
Browse more like this
Source: California Science & Engineering Fair public projects