It is notoriously hard to analyze sentiment in social media. It is written in a disorganized manner, using slang, and is
extremely dependent on implicit context. Classical baselines, including TF-IDF with Linear Support Vector Machines (SVM),
however, give a good starting point. Nevertheless, such approaches do not capture the underlying semantic meaning as they
deal with words individually. This paper compares traditional machine learning models with RoBERTa, transformer
architecture with the ability to capture the entire sentence context. Social media detain the real world are sloppy, with too fine
emotion tags and grossly skewed classes. In order to solve these problems, we designed a semantics-aware label consolidation
algorithm based on Sentence Transformers. We then balanced our training data by using SMOTE in the continuous embedding
space. Experiments with a multi-platform dataset with diverse data show that RoBERTa is significantly better than all classical
baselines in terms of accuracy, precision, recall, and F1-score. More to the point, the transformer is able to deal with sarcasm
and negation, which bag-of-words models often fail to do. This paper points to the need of contemporary sentiment analysis to
include context-sensitive embedding and specific preprocessing to be able to read online communication correctly
[1] K. Parmar, "Social Media Sentiments Analysis Dataset," Kaggle, 2023.
[2] N. V. Chawla, K. W. Bowyer, L. O. Hall, and W. P. Kegelmeyer, "SMOTE: Synthetic Minority Over-sampling
Technique," Journal of Artificial Intelligence Research, vol. 16, pp. 321–357, 2002.
[3] Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov, "RoBERTa:
A Robustly Optimized BERT Pretraining Approach," arXiv preprint arXiv:1907.11692, 2019.
[4] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, "BERT: Pre-training of Deep Bidirectional Transformers for
Language Understanding," in Proc. NAACL-HLT, Minneapolis, MN, USA, Jun. 2019, pp. 4171–4186.
[5] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, "Attention Is All
You Need," in Proc. NeurIPS, Long Beach, CA, USA, Dec. 2017, pp. 5998–6008.
[6] N. Reimers and I. Gurevych, "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks," in Proc.
EMNLP, Hong Kong, China, Nov. 2019, pp. 3982–3992.
[7] B. Pang and L. Lee, "Opinion Mining and Sentiment Analysis," Foundations and Trends in Information Retrieval, vol.
2, no. 1–2, pp. 1–135, 2008.
[8] S. M. Mohammad, "Sentiment Analysis: Detecting Valence, Emotions, and Other Affectual States from Text," in
Emotion Measurement, L. Tinianov, Ed. Elsevier, 2016, pp. 201–237.
[9] T. Joachims, "Text Categorization with Support Vector Machines: Learning with Many Relevant Features," in Proc.
ECML, Chemnitz, Germany, Apr. 1998, pp. 137–142.
[10] S. Bird, E. Loper, and E. Klein, Natural Language Processing with Python. Sebastopol, CA, USA: O'Reilly Media,
2009.
[11] F. Pedregosa et al., "Scikit-learn: Machine Learning in Python," Journal of Machine Learning Research, vol. 12, pp.
2825–2830, 2011.
[12] Y. Kim, "Convolutional Neural Networks for Sentence Classification," in Proc. EMNLP, Oct. 2014, pp. 1746–1751.
[13] T. Wolf et al., "HuggingFace's Transformers: State-of-the-Art Natural Language Processing," in Proc. EMNLP: System
Demonstrations, 2020, pp. 38–45.