Skip to main content
zenodoopen

Towards Generalization of Machine Learning Models: An Arabic Sentiment Analysis Dataset

<p>This data set consists of approximately 1.64 Million Arabic tweets (shared by their IDs) posted from 2009 to 2020, and their corresponding sentiment using a three-point classification system of Positive, Negative and Neutral/Mixed.&nbsp;No specific locations and/or keywords were specified throughout the data collection to obtain variation in the dialects and topics represented within the dataset. It is important to note that any biases in the proposed dataset in relation to the dialects and/or topics discussed were unintentional.</p> <p><strong>Please&nbsp;use the following citation if you use this data in a paper:</strong></p> <blockquote> <p>Abdaljalil, S., Hassanein, S., Mubarak, H., &amp; Abdelali, A. (2023). Towards Generalization of Machine Learning Models: A Case Study of Arabic Sentiment Analysis.&nbsp;<strong>Proceedings of the International AAAI Conference on Web and Social Media,&nbsp;17(1), 971-980.</strong></p> </blockquote> <p>&nbsp;</p> <p>&nbsp;</p>

ShareScore

40/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
20
Reuse readiness
8
Engagement
0

Topics