Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

34

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

34 results for “hate”

Learn how ShareScore rates datasets ↗
zenodo48/100

Salvaging the Internet Hate Machine: Using the discourse of extremist online subcultures to identify emergent extreme speech

<p>This dataset accompanies a paper submitted to the WebSci 20 conference.&nbsp;In this paper, we present a lexicon of &#39;extreme speech&#39; that may be used to detect hate speech and extreme speech on online platforms. We outline a cross-disciplinary research protocol through which this lexicon is initially extracted from a corpus of 3,335,265 posts from 4chan&#39;s /pol/ sub-forum using a hybrid method comprising word2vec modeling and subsequent snowballing of nearest neighbours of a small initial expert seed list of extreme language. The choice of corpus is significant, as 4chan is a space of rapid language innovation and obscure extreme vernacular, complicating generalised approaches. Our lexicon detects significantly more extreme posts within a corpus from a more mainstream platform (Reddit) than another popular lexicon, Hatebase, with similar accuracy. &nbsp;Our lexicon and the method of its creation thus provide a contribution to the study of the toxicity of online subcultures similar to 4chan, as well as more mainstream platforms. As we demonstrate, the lexicon allows for more effective detecting of extreme speech in these spaces. This method and the lexicon have further been made available through an open-source web tool for the study of online social platforms, 4CAT. The computational methods and lexicon on offer here can thus be used by a wide academic audience, fostering interdisciplinary approaches to the study of online hate and extreme speech.&nbsp;</p> <p>The dataset comprises the following items:</p> <ul> <li>The 4chan corpus from which the extreme speech lexicon was generated (posts from /pol/, 1 October 2019 - 1 November 2019)</li> <li>The Reddit corpus used to verify and test the lexicon (posts from the_donald, theredpill, politics and chapotraphouse, 1 October 2019 - 1 November 2019)</li> <li>The word2vec model from which the extreme speech lexicon was generated</li> <li>The extreme speech lexicon that was generated</li> </ul>

opencc-by-4.0Feb 2020View details →
zenodo40/100

TweetBLM: A Hate Speech Dataset and Analysis of BlackLivesMatter-related Microblogs on Twitter

<p>Collection of BLM related tweets and their corresponding labels of hate speech.</p>

opencc-by-4.0Aug 2020View details →
zenodo40/100

Hate Speech and Bias against Asians, Blacks, Jews, Latines, and Muslims: A Dataset for Machine Learning and Text Analytics

<h1>Institute for the Study of Contemporary Antisemitism (ISCA) at Indiana University Dataset on bias against Asians, Blacks, Jews, Latines, and Muslims&nbsp;</h1> <div> <h2>&nbsp;</h2> <h2>Description&nbsp;</h2> </div> <div> <p>The dataset is a product of a research project at Indiana University on biased messages on Twitter against ethnic and religious minorities. We scraped all live messages with the keywords "Asians, Blacks, Jews, Latinos, and Muslims" from the Twitter archive in 2020, 2021, and 2022.</p> <p>Random samples of 600 tweets were created for each keyword and year, including retweets. The samples were annotated in subsamples of 100 tweets by undergraduate students in Professor Gunther Jikeli's class 'Researching White Supremacism and Antisemitism on Social Media' in the fall of 2022 and 2023. A total of 120 students participated in 2022. They annotated datasets from 2020 and 2021. 134 students participated in 2023. They annotated datasets from the years 2021 and 2022. The annotation was done using the <a href="https://annotationportal.com/" target="_blank" rel="noreferrer noopener">Annotation Portal</a> (Jikeli, Soemer and Karali, 2024). The updated version of our portal, <a href="https://portal2.annotationportal.com/" target="_blank" rel="noreferrer noopener">AnnotHate</a>, is now publicly available. Each subsample was annotated by an average of 5.65 students per sample in 2022 and 8.32 students per sample in 2023, with a range of three to ten and three to thirteen students, respectively. Annotation included questions about bias and calling out bias.&nbsp;&nbsp;</p> </div> <div> <p>Annotators used a scale from 1 to 5 on the bias scale (confident not biased, probably not biased, don't know, probably biased, confident biased), using definitions of bias against each ethnic or religious group that can be found in the research reports from <a href="https://isca.indiana.edu/publication-research/social-media-project/Research-Report-BIAS-on-Twitter-against-Asians--Blacks-Jews-Latinos-Muslims-final-002.pdf" target="_blank" rel="noreferrer noopener">2022</a> and <a href="https://isca.indiana.edu/documents/BIAS%20Against%20Asian-Black-Hispanic-Jewish-and-%20Muslim-People%20on%20X-Twitter%20in%202021%20and%202022.pdf" target="_blank" rel="noreferrer noopener">2023</a>. If the annotators interpreted a message as biased according to the definition, they were instructed to choose the specific stereotype from the definition that was most applicable. Tweets that denounced bias against a minority were labeled as "calling out bias".&nbsp;&nbsp;&nbsp;</p> </div> <div> <p>The label was determined by a 75% majority vote. We classified &ldquo;probably biased&rdquo; and &ldquo;confident biased&rdquo; as biased, and &ldquo;confident not biased,&rdquo; &ldquo;probably not biased,&rdquo; and &ldquo;don't know&rdquo; as not biased.&nbsp;</p> </div> <div> <p>The stereotypes about the different minorities varied. About a third of all biased tweets were classified as general 'hate' towards the minority. The nature of specific stereotypes varied by group. Asians were blamed for the Covid-19 pandemic, alongside positive but harmful stereotypes about their perceived excessive privilege. Black people were associated with criminal activity and were subjected to views that portrayed them as inferior. Jews were depicted as wielding undue power and were collectively held accountable for the actions of the Israeli government. In addition, some tweets denied the Holocaust. Hispanic people/Latines faced accusations of being undocumented immigrants and "invaders," along with persistent stereotypes of them as lazy, unintelligent, or having too many children. Muslims were often collectively blamed for acts of terrorism and violence, particularly in discussions about Muslims in India.&nbsp;</p> </div> <div> <p>The annotation results from both cohorts (Class of 2022 and Class of 2023) will not be merged. They can be identified by the "cohort" column. While both cohorts (Class of 2022 and Class of 2023) annotated the same data from 2021,* their annotation results differ. The class of 2022 identified more tweets as biased for the keywords "Asians, Latinos, and Muslims" than the class of 2023, but nearly all of the tweets identified by the class of 2023 were also identified as biased by the class of 2022.&nbsp;&nbsp; The percentage of biased tweets with the keyword 'Blacks' remained nearly the same.&nbsp;&nbsp;&nbsp;</p> </div> <div> <p>*Due to a sampling error for the keyword "Jews" in 2021, the data are not identical between the two cohorts. The 2022 cohort annotated two samples for the keyword Jews, one from 2020 and the other from 2021, while the 2023 cohort annotated samples from 2021 and 2022.The 2021 sample for the keyword "Jews" that the 2022 cohort annotated was not representative. It has only 453 tweets from 2021 and 147 from the first eight months of 2022, and it includes some tweets from the query with the keyword "Israel". The 2021 sample for the keyword "Jews" that the 2023 cohort annotated was drawn proportionally for each trimester of 2021 for the keyword "Jews".&nbsp;</p> </div> <div> <h2>&nbsp;</h2> <h2>Content</h2> <h3>Cohort 2022&nbsp;</h3> </div> <div> <p>This dataset contains 5880 tweets that cover a wide range of topics common in conversations about Asians, Blacks, Jews, Latines, and Muslims. 357 tweets (6.1 %) are labeled as biased and 5523 (93.9 %) are labeled as not biased. 1365 tweets (23.2 %) are labeled as calling out or denouncing bias.&nbsp;&nbsp;</p> </div> <div> <p>1180 out of 5880 tweets (20.1 %) contain the keyword "Asians," 590 were posted in 2020 and 590 in 2021. 39 tweets (3.3 %) are biased against Asian people. 370 tweets (31,4 %) call out bias against Asians.&nbsp;&nbsp;</p> </div> <div> <p>1160 out of 5880 tweets (19.7%) contain the keyword "Blacks," 578 were posted in 2020 and 582 in 2021. 101 tweets (8.7 %) are biased against Black people. 334 tweets (28.8 %) call out bias against Blacks.&nbsp;&nbsp;</p> </div> <div> <p>1189 out of 5880 tweets (20.2 %) contain the keyword "Jews," 592 were posted in 2020, 451 in 2021, and &ndash;&ndash;as mentioned above&ndash;&ndash;146 tweets from 2022. 83 tweets (7 %) are biased against Jewish people. 220 tweets (18.5 %) call out bias against Jews.&nbsp;</p> </div> <div> <p>1169 out of 5880 tweets (19.9 %) contain the keyword "Latinos," 584 were posted in 2020 and 585 in 2021. 29 tweets (2.5 %) are biased against Latines. 181 tweets (15.5 %) call out bias against Latines.&nbsp;&nbsp;</p> </div> <div> <p>1182 out of 5880 tweets (20.1 %) contain the keyword "Muslims," 593 were posted in 2020 and 589 in 2021. 105 tweets (8.9 %) are biased against Muslims. 260 tweets (22 %) call out bias against Muslims.&nbsp;&nbsp;</p> </div> <div> <h3>Cohort 2023&nbsp;</h3> </div> <div> <p>The dataset contains 5363 tweets with the keywords &ldquo;Asians, Blacks, Jews, Latinos and Muslims&rdquo; from 2021 and 2022. 261 tweets (4.9 %) are labeled as biased, and 5102 tweets (95.1 %) were labeled as not biased. 975 tweets (18.1 %) were labeled as calling out or denouncing bias.&nbsp;</p> </div> <div> <p>1068 out of 5363 tweets (19.9 %) contain the keyword "Asians," 559 were posted in 2021 and 509 in 2022. 42 tweets (3.9 %) are biased against Asian people. 280 tweets (26.2 %) call out bias against Asians.&nbsp;&nbsp;</p> </div> <div> <p>1130 out of 5363 tweets (21.1 %) contain the keyword "Blacks," 586 were posted in 2021 and 544 in 2022. 76 tweets (6.7 %) are biased against Black people. 146 tweets (12.9 %) call out bias against Blacks.&nbsp;&nbsp;</p> </div> <div> <p>971 out of 5363 tweets (18.1 %) contain the keyword "Jews," 460 were posted in 2021 and 511 in 2022. 49 tweets (5 %) are biased against Jewish people. 201 tweets (20.7 %) call out bias against Jews.&nbsp;</p> </div> <div> <p>1072 out of 5363 tweets (19.9 %) contain the keyword "Latinos," 583 were posted in 2021 and 489 in 2022. 32 tweets (2.9 %) are biased against Latines. 108 tweets (10.1 %) call out bias against Latines.&nbsp;&nbsp;</p> </div> <div> <p>1122 out of 5363 tweets (20.9 %) contain the keyword "Muslims," 576 were posted in 2021 and 546 in 2022. 62 tweets (5.5 %) are biased against Muslims. 240 tweets (21.3 %) call out bias against Muslims.&nbsp;</p> </div> <div> <h2>&nbsp;</h2> <h2>File Description</h2> </div> <div> <p>The dataset is provided in a csv file format, with each row representing a single message, including replies, quotes, and retweets. The file contains the following columns:&nbsp;&nbsp;</p> <p>'TweetID': Represents the tweet ID.&nbsp;&nbsp;&nbsp;</p> </div> <div> <p>'Username': Represents the username who published the tweet (if it is a retweet, it will be the user who retweetet the original tweet.&nbsp;&nbsp;&nbsp;&nbsp;</p> </div> <div> <p>'Text': Represents the full text of the tweet (not pre-processed).&nbsp;&nbsp;</p> </div> <div> <p>'CreateDate': Represents the date the tweet was created.&nbsp;&nbsp;&nbsp;&nbsp;</p> </div> <div> <p>'Biased': Represents the labeled by our annotators if the tweet is biased (1) or not (0).&nbsp;&nbsp;</p> </div> <div> <p>'Calling_Out': Represents the label by our annotators if the tweet is calling out bias against minority groups (1) or not (0).&nbsp;&nbsp;</p> </div> <div> <p>'Keyword': Represents the keyword that was used in the query. The keyword can be in the text, including mentioned names, or the username.&nbsp;&nbsp;&nbsp;&nbsp;</p> </div> <div> <p>&nbsp;&lsquo;Cohort&rsquo;: Represents the year the data was annotated (class of 2022 or class of 2023)&nbsp;</p> </div> <div> <h2>&nbsp;</h2> <h2>Acknowledgements&nbsp; &nbsp;</h2> </div> <div> <p>We are grateful for the technical collaboration with Indiana University's Observatory on Social Media (OSoMe). We thank all class participants for the annotations and contributions, including Kate Baba, Eleni Ballis, Garrett Banuelos, Savannah Benjamin, Luke Bianco, Zoe Bogan, Elisha S. Breton, Aidan Calderaro, Anaye Caldron, Olivia Cozzi, Daj Crisler, Jenna Eidson, Ella Fanning, Victoria Ford, Jess Gruettner, Ronan Hancock, Isabel Hawes, Brennan Hensler, Kyra Horton, Maxwell Idczak, Sanjana Iyer, Jacob Joffe, Katie Johnson, Allison Jones, Kassidy Keltner, Sophia Knoll, Jillian Kolesky, Emily Lowrey, Rachael Morara, Benjamin Nadolne, Rachel Neglia, Seungmin Oh, Kirsten Pecsenye, Sophia Perkovich, Joey Philpott, Katelin Ray, Kaleb Samuels, Chloe Sherman, Rachel Weber, Molly Winkeljohn, Ally Wolfgang, Rowan Wolke, Michael Wong, Jane Woods, Kaleb Woodworth, Aurora Young, Sydney Allen, Hundre Askie, Norah Bardol, Olivia Baren, Samuel Barth, Emma Bender, Noam Biron, Kendyl Bond, Graham Brumley, Kennedi Bruns, Leah Burger, Hannah Busche, Morgan Butrum-Griffith, Zoe Catlin, Angeli Cauley, Nathalya Chavez Medrano, Mia Cooper, Suhani Desai, Isabella Flick, Samantha Garcez, Isabella Grady, Macy Hutchinson, Sarah Kirkman, Ella Leitner, Elle Marquardt, Madison Moss, Ethan Nixdorf, Reya Patel, Mickey Racenstein, Kennedy Rehklau, Grace Roggeman, Jack Rossell, Madeline Rubin, Fernando Sanchez, Hayden Sawyer, Diego Scheker, Lily Schwecke, Brooke Scott, Megan Scott, Samantha Secchi, Jolie Segal, Katherine Smith, Constantine Stefanidis, Cami Stetler, Madisyn West, Alivia Yusefzadeh, Tayssir Aminou, Karen Fecht, Luciana Orrego-Hoyos, Hannah Pickett, and Sophia Tracy.&nbsp;</p> </div> <div> <p>This work used Jetstream2 at Indiana University through allocation HUM200003 from the Advanced Cyberinfrastructure Coordination Ecosystem: Services &amp; Support (ACCESS) program, which is supported by National Science Foundation grants #2138259, #2138286, #2138307, #2137603, and #2138296.&nbsp;</p> </div> <div> <p>&nbsp;</p> </div>

opencc-by-4.0Mar 2023View details →
zenodo40/100

Algerian Dialect Dataset Targeted Hate Speech, Offensive Language and Cyberbullying

<p>Algerian Dialect Dataset Targeted Hate Speech, Offensive Language and Cyberbullying.</p> <p>&nbsp;</p> <p>* To cite this dataset refer to&nbsp;<a href="http://dx.doi.org/10.12785/ijcds/130177" target="_blank" rel="nofollow noopener">http://dx.doi.org/10.12785/ijcds/130177</a><br>Mazari, A. C., &amp; Kheddar, H. (2023). "Deep Learning-based Analysis of Algerian Dialect Dataset Targeted Hate Speech, Offensive Language and Cyberbullying." IJCDS, 13(1).</p> <p>&nbsp;</p> <div> <p>* Due to the nature of this Dataset, comments contain offensiveness and hate speech. This does not reflect author values, however the aim is to providing a resource to help in detecting and preventing spread of such harmful content.</p> </div> <div> <h3>Features</h3> <ul> <li>Algerian Dialect</li> <li>Cyberbullying</li> <li>Hate speech</li> <li>Offensive Language</li> <li>Dialect Dataset</li> </ul> </div>

opencc-by-4.0Apr 2024View details →
zenodo40/100

On the Effectiveness of Text and Image Embeddings in Multimodal Hate Speech Detection

<p>Additional resources for the paper:</p> <h3><strong><a href="https://ieeexplore.ieee.org/abstract/document/10826088">On the Effectiveness of Text and Image Embeddings in Multimodal Hate Speech Detection.</a></strong></h3> <p>Lewis, N., Cavalcante, C. C., Boukouvalas, Z., &amp; Corizzo, R.</p> <p><em>2024 IEEE International Conference on Big Data (BigData)</em> (pp. 3277-3281). IEEE.</p> <pre>&nbsp;</pre> <p>&nbsp;</p> <p>MMHS150K [1] is a manually labeled multimodal dataset that contains $150000$ tweets with two modalities: text, and &nbsp;corresponding image. Tweets are collected from September 2018 until February 2019 and are labeled according to different types of hate speech: no attacks to any community, racist, sexist, homophobic, religion-based attacks, or attacks to other communities.&nbsp;</p> <p>We extract vector embeddings leveraging different text (BERT, OpenAI) and image (ResNet, PVT, ViT) modele backbones and assess their effectiveness in the hate speech detection task.</p> <p>&nbsp;</p> <h2>Citation:</h2> <pre>@inproceedings{lewis2024effectiveness, title={On the Effectiveness of Text and Image Embeddings in Multimodal Hate Speech Detection}, author={Lewis, Nora and Cavalcante, Charles C and Boukouvalas, Zois and Corizzo, Roberto}, booktitle={2024 IEEE International Conference on Big Data (BigData)}, pages={3277--3281}, year={2024}, organization={IEEE} }</pre>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Detecting weak and strong Islamophobic hate speech on social media

<p>Data, code and annotation guidelines for our publication, &#39;Detecting weak and strong Islamophobic hate speech on social media&#39; (2019).</p>

opencc-by-4.0Sep 2019View details →
zenodo40/100

Anderson Police Department Hate Crime Demographics 2023

<p>This dataset contains demographic information related to reported hate crimes within the jurisdiction of the Anderson Police Department. The data includes details on both victims and alleged perpetrators, with demographic variables such as age, gender, and race/ethnicity. The types of hate crimes covered in the dataset are based on classifications in accordance with relevant local and federal hate crime definitions.</p> <p>The dataset was obtained through a Public Records Act request and covers the time period from January 1st 2023 through December 31st 2023. This agency had no Hate Crimes to report in 2022. It was provided in MS Word format, where each row represents a unique hate crime incident and the columns capture demographic and other related variables.</p>

opencc-by-4.0Jul 2024View details →
zenodo40/100

Amharic Hate Speech Detection Dataset

<p>Amharic Hate Speech Detection Dataset V1</p> <p>To contribute for the research and development of hate speech detection in Amharic language from social media, we are glad to release our hate speech dataset we prepared from the Ethiopian Broadcasting Corporation (EBC) Facebook page (<a href="https://www.facebook.com/EBCzena">https://www.facebook.com/EBCzena</a>), and some chosen Facebook page (<a href="https://www.facebook.com/604407519910492">https://www.facebook.com/604407519910492</a>) that we found potential hateful comments.</p> <p>We extracted comments/posts pertaining to race, religion, and ethnicity using the Facepager API, resulting in a set of 30,000 comments between April 15, 2019 and December 15, 2019. A total of 5,000 comments/posts were chosen at random for annotation. Three annotators (two candidate PhD. in Linguistics and one MSc. in Law) manually annotated the selected samples as &ldquo;<strong>Hate</strong>&rdquo; or &ldquo;<strong>not</strong>-<strong>Hate</strong>&rdquo; resulting 2,000 (1000 hate and 1000 non-hate) labeled comments because of majority vote among the annotators.</p> <p>For the labeling procedure, the annotators used Ethiopian government&rsquo;s hate speech and misinformation prevention and suppression proclamation <a href="https://www.accessnow.org/cms/assets/uploads/2020/05/Hate-Speech-and-Disinformation-Prevention-and-Suppression-Proclamation.pdf">https://www.accessnow.org/cms/assets/uploads/2020/05/Hate-Speech-and-Disinformation-Prevention-and-Suppression-Proclamation.pdf</a>, as well as our definition of hate speech and the hate speech characterization lists proposed in Fino (2020) <a href="https://doi.org/10.1093/jicj/mqaa023">https://doi.org/10.1093/jicj/mqaa023</a>) were provided to the annotators.</p> <p>Accordingly, a speech is labeled as &ldquo;<strong>Hate</strong>&rdquo; when:</p> <ul> <li>&ldquo;the speech targets a group or individual as a member of a group (ethnicity, race, religion)&rdquo;</li> <li>&ldquo;the speech content in the message expresses hatred&rdquo;</li> <li>&ldquo;the speech causes a harm&rdquo;</li> <li>&ldquo;the speaker intends harm or bad activity&rdquo;</li> <li>&ldquo;the speech incites bad actions&rdquo;</li> <li>&ldquo;the speech is either public and directed at a member of the group&rdquo;</li> <li>&ldquo;the context makes violent response possible&rdquo;</li> </ul>

opencc-by-4.0Jun 2021View details →
zenodo36/100

Mpox Narrative on Instagram: A Labeled Multilingual Dataset of Instagram Posts on Mpox for Sentiment, Hate Speech, and Anxiety Analysis

<p><strong>Please cite the following paper when using this dataset</strong>:</p> <p>N. Thakur, &ldquo;Mpox narrative on Instagram: A labeled multilingual dataset of Instagram posts on mpox for sentiment, hate speech, and anxiety analysis,&rdquo; arXiv [cs.LG], 2024, URL: https://arxiv.org/abs/2409.05292</p> <p><strong>Abstract</strong></p> <p>The world is currently experiencing an outbreak of mpox, which has been declared a Public Health Emergency of International Concern by WHO. During recent virus outbreaks, social media platforms have played a crucial role in keeping the global population informed and updated regarding various aspects of the outbreaks. As a result, in the last few years, researchers from different disciplines have focused on the development of social media datasets focusing on different virus outbreaks. No prior work in this field has focused on the development of a dataset of Instagram posts about the mpox outbreak. The work presented in this paper (stated above) aims to address this research gap. It presents this <strong>multilingual dataset of</strong>&nbsp;<strong>60,127 Instagram posts</strong> about mpox, published between <strong>July 23, 2022, and September 5, 2024</strong>. This dataset contains Instagram posts about mpox in <strong>52 languages</strong>. For each of these posts, the Post ID, Post Description, Date of publication, language, and translated version of the post (translation to English was performed using the Google Translate API) are presented as separate attributes in the dataset.</p> <p>After developing this dataset, sentiment analysis, hate speech detection, and anxiety or stress detection were also performed. This process included classifying each post into</p> <ul> <li>one of the fine-grain sentiment classes, i.e., <strong>fear, surprise, joy, sadness, anger, disgust, or neutral</strong>,&nbsp;</li> <li><strong>hate or not hate</strong></li> <li><strong>anxiety/stress detected or no anxiety/stress detected</strong>.</li> </ul> <p>These results are presented as separate attributes in the dataset for the training and testing of machine learning algorithms for sentiment, hate speech, and anxiety or stress detection, as well as for other applications.&nbsp;</p> <p><strong>The 52 distinct languages in which Instagram posts are present in the dataset&nbsp;</strong><strong>are&nbsp;</strong>English, Portuguese, Indonesian, Spanish, Korean, French, Hindi, Finnish, Turkish, Italian, German, Tamil, Urdu, Thai, Arabic, Persian, Tagalog, Dutch, Catalan, Bengali, Marathi, Malayalam, Swahili, Afrikaans, Panjabi, Gujarati, Somali, Lithuanian, Norwegian, Estonian, Swedish, Telugu, Russian, Danish, Slovak, Japanese, Kannada, Polish, Vietnamese, Hebrew, Romanian, Nepali, Czech, Modern Greek, Albanian, Croatian, Slovenian, Bulgarian, Ukrainian, Welsh, Hungarian, and Latvian.&nbsp;</p> <p>The following table represents the data description for this dataset</p> <table> <tbody> <tr> <td> <p><strong>Attribute Name</strong></p> </td> <td> <p><strong>Attribute Description</strong></p> </td> </tr> <tr> <td> <p>Post ID</p> </td> <td> <p>Unique ID of each Instagram post</p> </td> </tr> <tr> <td> <p>Post Description</p> </td> <td> <p>Complete description of each post in the language in which it was originally published</p> </td> </tr> <tr> <td> <p>Date</p> </td> <td> <p>Date of publication in MM/DD/YYYY format</p> </td> </tr> <tr> <td> <p>Language</p> </td> <td> <p>Language of the post as detected using the Google Translate API</p> </td> </tr> <tr> <td> <p>Translated Post Description</p> </td> <td> <p>Translated version of the post description. All posts which were not in English were translated into English using the Google Translate API. No language translation was performed for English posts.</p> </td> </tr> <tr> <td> <p>Sentiment</p> </td> <td> <p>Results of sentiment analysis (using translated Post Description) where each post was classified into one of the sentiment classes: fear, surprise, joy, sadness, anger, disgust, and neutral</p> </td> </tr> <tr> <td> <p>Hate</p> </td> <td> <p>Results of hate speech detection (using translated Post Description) where each post was classified as hate or not hate</p> </td> </tr> <tr> <td> <p>Anxiety or Stress</p> </td> <td> <p>Results of anxiety or stress detection (using translated Post Description) where each post was classified as stress/anxiety detected or no stress/anxiety detected.</p> </td> </tr> </tbody> </table>

opencc-by-4.0Dec 2023View details →
zenodo36/100

Indonesian Foreign Policy towards Iran: Shia Hate speech on X social media with LSTM and SVM Analysis

<p>This table and figure are integral components of research on Shia hate speech on social media X, a critical issue in the context of identity politics in Indonesia and globally. This study is of paramount importance as it delves into the identity politics often exploited by politicians in Indonesia and around the world. The Shia community's support for President Jokowi in the first and second stages of the Election was met with hate speech from the opposition group. The study further investigates whether this Shia hate speech is linked to the government's policy towards Iran, a country known for its Shia ideology. The study is presented in three parts:<br>1. Table detailing the sentiment analysis process and results, which were conducted using advanced machine learning techniques such as SVM and LSTM. This approach significantly enhances the accuracy and reliability of the study's findings.<br>2. Figure in the form of a graph related to the study results and the results of comments from the Indonesian public about Shia.<br>3. This research is backed by a comprehensive dataset comprising public comments from Indonesia on Shia. This extensive data collection ensures the study's conclusions are thorough and reliable.</p> <p>4. Processing of machine learning</p>

opencc-by-4.0Oct 2024View details →
zenodo32/100

League of Legends and hate speech: a corpus for comments in Twitch.tv

<p>League of Legends (LOL) is the most popular game on PC, drawing 8 million concurrent players. A common activity of gamers, besides playing games, is to watch other players presenting tips and tricks. Streaming platforms allow some players to show gameplays and live games. <a href="https://www.twitch.tv/">Twitch.tv</a> is the world&acute;s leading live streaming platform.&nbsp;</p> <p>Considering that hate speech is a ubiquitous problem in online gaming, we collected &nbsp;985,766 comments from five videos of the top 10 &nbsp;LOL streamers in Twitch.tv platform.&nbsp;</p> <p>The dataset is freely available in a single file, ensembling all videos/players; and divided by players as well.&nbsp;</p> <p>These comments are a rich data source for opinion mining, sentiment analysis, topic modeling, and hate speech detection (including sexism and racism).</p> <ul> </ul>

opencc-by-4.0Mar 2020View details →
zenodo32/100

Marathi Hate Comments

<p>The dataset consists of a CSV file of comments in the Marathi language. It includes comments from various social media platforms such as Twitter, Facebook, Instagram, and YouTube. The dataset has columns with column names as id, hashtag, text, date, label, and platform respectively. The label column has values either HOF, indicating offensive text, or NOT, indicating a non-offensive text. The number of comments in the file is around 2103</p>

opencc-by-nc-sa-4.0Mar 2022View details →
zenodo32/100

Replication package for: Paying Them to Hate US: The Effect of U.S. Military Aid on Anti-American Terrorism, 1968-2018

<p>This replication package contains the data and the code to replicate the tables of the paper "Paying Them to Hate US: The Effect of U.S. Military Aid on Anti-American Terrorism, 1968-2018", conditionally accepted for publication in The Economic Journal.</p>

opencc-by-4.0Apr 2024View details →
zenodo32/100

HaterNet a system for detecting and analyzing hate speech in Twitter

<p>This dataset consists of&nbsp; two corpuses used in the paper &quot;Detecting and analyzing hate speech in Twitter: HaterNet a system in the Spanish prevention of hate crime office&quot;. A first one based on tweets collected at different random dates between February 2017 and December 2017 with a final size of 2 million tweets. A second one with&nbsp;6,000 tweets labeled as described in the paper as hate containing or not.</p>

opencc-by-4.0Mar 2019View details →
zenodo32/100

HOCON34k: A Corpus of Hate speech in Online Comments from German Newspapers

<p>We have compiled a dataset containing 34,223 comments in German, authored by users from online-platforms associated with public discourse in German newspapers. Each comment was annotated for hate speech and the adequacy of contextual information by a group of 29 volunteers, using a binary annotation approach. The inter-rater reliability for hate speech is 0.4428 across all annotators and increases to 0.6078 when considering an optimized subset of 12 annotators, as measured by Fleiss&rsquo; Kappa. Additionally, we present a baseline text classification using BERT, achieving an MCC-score up to 0.32 and an F2-score up to 0.64 in our initial experiment on this new corpus. The data set, named HOCON34k, comprising German hate speech comments from newspapers, is publicly available for research purposes.</p>

opencc-by-4.0Dec 2024View details →
zenodo32/100

Dataset for: From criticism to anger and hate: The vulgarisation of digital press criticism on news outlets' Facebook page

<p>Dataset of comments for publication From criticism to anger and hate: The vulgarisation of digital press criticism on news outlets&rsquo; Facebook page.</p>

opencc-by-4.0Nov 2022View details →
dryad28/100

Data from: Why hate the good guy? Antisocial punishment of high cooperators is greater when people compete to be chosen

When choosing social partners, people prefer good cooperators (all else equal). Given this preference, anyone wishing to be chosen can either increase their own cooperation to become more desirable, or suppress others' cooperation to make them less desirable. Previous research shows that very cooperative people sometimes get punished ("antisocial punishment") or criticized ("do-gooder derogation") in many cultures. Here we use a public goods game with punishment to test whether antisocial punishment is used as a means of competing to be chosen by suppressing others' cooperation. As predicted, there was more antisocial punishment when participants were competing to be chosen for a subsequent cooperative task (a Trust Game) than without a subsequent task. This difference in antisocial punishment cannot be explained by differences in contributions, moralistic punishment, or confusion. This suggests that antisocial punishment is a social strategy that low cooperators use to avoid looking bad when high cooperators escalate cooperation.

opencc-zeroDec 2016View details →
zenodo28/100

Hate Speech

Open the record for dataset details and reuse information.

opencc-by-4.0Jul 2024View details →
zenodo28/100

Figure 1 in Xenophobia, Radicalism, and Hate Crime in Europe Annual Report

Figure 1. – Costello graph (modified by Amundsen et al., 1996): relationship between frequency of occurrence (%F) of prey items and prey-specific abundance (Pi), expressed as number (left) and weight (right), respectively, in the diet of Oblada melanura collected in the Strait of Sicily. The graph background shows the explanatory Costello diagram and its interpretation of feeding strategy (BPC = betweenphenotype component; WPC = within-phenotype component).

opencc-by-4.0Dec 2018View details →
zenodo28/100

Unsafe Diffusion: On the Generation of Unsafe Images and Hateful Memes From Text-To-Image Models

<h2>[Update] Looking for a larger unsafe image dataset? We publish a new dataset named UnsafeBench on Hugging Face. Take a look at&nbsp;<a href="https://huggingface.co/datasets/yiting/UnsafeBench">here</a>!</h2> <p>This dataset used in the paper&nbsp; <a href="https://arxiv.org/pdf/2305.13873.pdf">https://arxiv.org/pdf/2305.13873.pdf</a>&nbsp;contains four prompt sets and one image set.</p> <p>The four prompt sets were used to query Text-to-Image models and generate images for safety assessment. These sets include three&nbsp;harmful prompt sets&nbsp;and one harmless prompt&nbsp;set. The harmful prompts originate from different sources and contain&nbsp;various unsafe concepts, such as sexually explicit, violent, disturbing, hateful, and political content.</p> <p><strong>Prompt Sets</strong>:</p> <ul> <li>4chan Prompts: Harmful</li> <li>Lexica Prompts: Harmful</li> <li>Template Prompts: Harmful</li> <li>COCO Prompts: Harmless</li> </ul> <p><strong>Image Dataset</strong>:</p> <p>This dataset consists of 800 images, which were randomly selected from all the generated images from Text-to-Image models.</p> <ul> <li>Safe: 580 images</li> <li>Sexually Explicit: 48 images</li> <li>Violent: 45 images</li> <li>Disturbing: 68 images</li> <li>Hateful: 35 images</li> <li>Political: 50 images</li> </ul>

openAug 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record