Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

25

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

25 results for “Hate speech”

Learn how ShareScore rates datasets ↗
zenodo48/100

Salvaging the Internet Hate Machine: Using the discourse of extremist online subcultures to identify emergent extreme speech

<p>This dataset accompanies a paper submitted to the WebSci 20 conference.&nbsp;In this paper, we present a lexicon of &#39;extreme speech&#39; that may be used to detect hate speech and extreme speech on online platforms. We outline a cross-disciplinary research protocol through which this lexicon is initially extracted from a corpus of 3,335,265 posts from 4chan&#39;s /pol/ sub-forum using a hybrid method comprising word2vec modeling and subsequent snowballing of nearest neighbours of a small initial expert seed list of extreme language. The choice of corpus is significant, as 4chan is a space of rapid language innovation and obscure extreme vernacular, complicating generalised approaches. Our lexicon detects significantly more extreme posts within a corpus from a more mainstream platform (Reddit) than another popular lexicon, Hatebase, with similar accuracy. &nbsp;Our lexicon and the method of its creation thus provide a contribution to the study of the toxicity of online subcultures similar to 4chan, as well as more mainstream platforms. As we demonstrate, the lexicon allows for more effective detecting of extreme speech in these spaces. This method and the lexicon have further been made available through an open-source web tool for the study of online social platforms, 4CAT. The computational methods and lexicon on offer here can thus be used by a wide academic audience, fostering interdisciplinary approaches to the study of online hate and extreme speech.&nbsp;</p> <p>The dataset comprises the following items:</p> <ul> <li>The 4chan corpus from which the extreme speech lexicon was generated (posts from /pol/, 1 October 2019 - 1 November 2019)</li> <li>The Reddit corpus used to verify and test the lexicon (posts from the_donald, theredpill, politics and chapotraphouse, 1 October 2019 - 1 November 2019)</li> <li>The word2vec model from which the extreme speech lexicon was generated</li> <li>The extreme speech lexicon that was generated</li> </ul>

opencc-by-4.0Feb 2020View details →
zenodo40/100

TweetBLM: A Hate Speech Dataset and Analysis of BlackLivesMatter-related Microblogs on Twitter

<p>Collection of BLM related tweets and their corresponding labels of hate speech.</p>

opencc-by-4.0Aug 2020View details →
zenodo40/100

Hate Speech and Bias against Asians, Blacks, Jews, Latines, and Muslims: A Dataset for Machine Learning and Text Analytics

<h1>Institute for the Study of Contemporary Antisemitism (ISCA) at Indiana University Dataset on bias against Asians, Blacks, Jews, Latines, and Muslims&nbsp;</h1> <div> <h2>&nbsp;</h2> <h2>Description&nbsp;</h2> </div> <div> <p>The dataset is a product of a research project at Indiana University on biased messages on Twitter against ethnic and religious minorities. We scraped all live messages with the keywords "Asians, Blacks, Jews, Latinos, and Muslims" from the Twitter archive in 2020, 2021, and 2022.</p> <p>Random samples of 600 tweets were created for each keyword and year, including retweets. The samples were annotated in subsamples of 100 tweets by undergraduate students in Professor Gunther Jikeli's class 'Researching White Supremacism and Antisemitism on Social Media' in the fall of 2022 and 2023. A total of 120 students participated in 2022. They annotated datasets from 2020 and 2021. 134 students participated in 2023. They annotated datasets from the years 2021 and 2022. The annotation was done using the <a href="https://annotationportal.com/" target="_blank" rel="noreferrer noopener">Annotation Portal</a> (Jikeli, Soemer and Karali, 2024). The updated version of our portal, <a href="https://portal2.annotationportal.com/" target="_blank" rel="noreferrer noopener">AnnotHate</a>, is now publicly available. Each subsample was annotated by an average of 5.65 students per sample in 2022 and 8.32 students per sample in 2023, with a range of three to ten and three to thirteen students, respectively. Annotation included questions about bias and calling out bias.&nbsp;&nbsp;</p> </div> <div> <p>Annotators used a scale from 1 to 5 on the bias scale (confident not biased, probably not biased, don't know, probably biased, confident biased), using definitions of bias against each ethnic or religious group that can be found in the research reports from <a href="https://isca.indiana.edu/publication-research/social-media-project/Research-Report-BIAS-on-Twitter-against-Asians--Blacks-Jews-Latinos-Muslims-final-002.pdf" target="_blank" rel="noreferrer noopener">2022</a> and <a href="https://isca.indiana.edu/documents/BIAS%20Against%20Asian-Black-Hispanic-Jewish-and-%20Muslim-People%20on%20X-Twitter%20in%202021%20and%202022.pdf" target="_blank" rel="noreferrer noopener">2023</a>. If the annotators interpreted a message as biased according to the definition, they were instructed to choose the specific stereotype from the definition that was most applicable. Tweets that denounced bias against a minority were labeled as "calling out bias".&nbsp;&nbsp;&nbsp;</p> </div> <div> <p>The label was determined by a 75% majority vote. We classified &ldquo;probably biased&rdquo; and &ldquo;confident biased&rdquo; as biased, and &ldquo;confident not biased,&rdquo; &ldquo;probably not biased,&rdquo; and &ldquo;don't know&rdquo; as not biased.&nbsp;</p> </div> <div> <p>The stereotypes about the different minorities varied. About a third of all biased tweets were classified as general 'hate' towards the minority. The nature of specific stereotypes varied by group. Asians were blamed for the Covid-19 pandemic, alongside positive but harmful stereotypes about their perceived excessive privilege. Black people were associated with criminal activity and were subjected to views that portrayed them as inferior. Jews were depicted as wielding undue power and were collectively held accountable for the actions of the Israeli government. In addition, some tweets denied the Holocaust. Hispanic people/Latines faced accusations of being undocumented immigrants and "invaders," along with persistent stereotypes of them as lazy, unintelligent, or having too many children. Muslims were often collectively blamed for acts of terrorism and violence, particularly in discussions about Muslims in India.&nbsp;</p> </div> <div> <p>The annotation results from both cohorts (Class of 2022 and Class of 2023) will not be merged. They can be identified by the "cohort" column. While both cohorts (Class of 2022 and Class of 2023) annotated the same data from 2021,* their annotation results differ. The class of 2022 identified more tweets as biased for the keywords "Asians, Latinos, and Muslims" than the class of 2023, but nearly all of the tweets identified by the class of 2023 were also identified as biased by the class of 2022.&nbsp;&nbsp; The percentage of biased tweets with the keyword 'Blacks' remained nearly the same.&nbsp;&nbsp;&nbsp;</p> </div> <div> <p>*Due to a sampling error for the keyword "Jews" in 2021, the data are not identical between the two cohorts. The 2022 cohort annotated two samples for the keyword Jews, one from 2020 and the other from 2021, while the 2023 cohort annotated samples from 2021 and 2022.The 2021 sample for the keyword "Jews" that the 2022 cohort annotated was not representative. It has only 453 tweets from 2021 and 147 from the first eight months of 2022, and it includes some tweets from the query with the keyword "Israel". The 2021 sample for the keyword "Jews" that the 2023 cohort annotated was drawn proportionally for each trimester of 2021 for the keyword "Jews".&nbsp;</p> </div> <div> <h2>&nbsp;</h2> <h2>Content</h2> <h3>Cohort 2022&nbsp;</h3> </div> <div> <p>This dataset contains 5880 tweets that cover a wide range of topics common in conversations about Asians, Blacks, Jews, Latines, and Muslims. 357 tweets (6.1 %) are labeled as biased and 5523 (93.9 %) are labeled as not biased. 1365 tweets (23.2 %) are labeled as calling out or denouncing bias.&nbsp;&nbsp;</p> </div> <div> <p>1180 out of 5880 tweets (20.1 %) contain the keyword "Asians," 590 were posted in 2020 and 590 in 2021. 39 tweets (3.3 %) are biased against Asian people. 370 tweets (31,4 %) call out bias against Asians.&nbsp;&nbsp;</p> </div> <div> <p>1160 out of 5880 tweets (19.7%) contain the keyword "Blacks," 578 were posted in 2020 and 582 in 2021. 101 tweets (8.7 %) are biased against Black people. 334 tweets (28.8 %) call out bias against Blacks.&nbsp;&nbsp;</p> </div> <div> <p>1189 out of 5880 tweets (20.2 %) contain the keyword "Jews," 592 were posted in 2020, 451 in 2021, and &ndash;&ndash;as mentioned above&ndash;&ndash;146 tweets from 2022. 83 tweets (7 %) are biased against Jewish people. 220 tweets (18.5 %) call out bias against Jews.&nbsp;</p> </div> <div> <p>1169 out of 5880 tweets (19.9 %) contain the keyword "Latinos," 584 were posted in 2020 and 585 in 2021. 29 tweets (2.5 %) are biased against Latines. 181 tweets (15.5 %) call out bias against Latines.&nbsp;&nbsp;</p> </div> <div> <p>1182 out of 5880 tweets (20.1 %) contain the keyword "Muslims," 593 were posted in 2020 and 589 in 2021. 105 tweets (8.9 %) are biased against Muslims. 260 tweets (22 %) call out bias against Muslims.&nbsp;&nbsp;</p> </div> <div> <h3>Cohort 2023&nbsp;</h3> </div> <div> <p>The dataset contains 5363 tweets with the keywords &ldquo;Asians, Blacks, Jews, Latinos and Muslims&rdquo; from 2021 and 2022. 261 tweets (4.9 %) are labeled as biased, and 5102 tweets (95.1 %) were labeled as not biased. 975 tweets (18.1 %) were labeled as calling out or denouncing bias.&nbsp;</p> </div> <div> <p>1068 out of 5363 tweets (19.9 %) contain the keyword "Asians," 559 were posted in 2021 and 509 in 2022. 42 tweets (3.9 %) are biased against Asian people. 280 tweets (26.2 %) call out bias against Asians.&nbsp;&nbsp;</p> </div> <div> <p>1130 out of 5363 tweets (21.1 %) contain the keyword "Blacks," 586 were posted in 2021 and 544 in 2022. 76 tweets (6.7 %) are biased against Black people. 146 tweets (12.9 %) call out bias against Blacks.&nbsp;&nbsp;</p> </div> <div> <p>971 out of 5363 tweets (18.1 %) contain the keyword "Jews," 460 were posted in 2021 and 511 in 2022. 49 tweets (5 %) are biased against Jewish people. 201 tweets (20.7 %) call out bias against Jews.&nbsp;</p> </div> <div> <p>1072 out of 5363 tweets (19.9 %) contain the keyword "Latinos," 583 were posted in 2021 and 489 in 2022. 32 tweets (2.9 %) are biased against Latines. 108 tweets (10.1 %) call out bias against Latines.&nbsp;&nbsp;</p> </div> <div> <p>1122 out of 5363 tweets (20.9 %) contain the keyword "Muslims," 576 were posted in 2021 and 546 in 2022. 62 tweets (5.5 %) are biased against Muslims. 240 tweets (21.3 %) call out bias against Muslims.&nbsp;</p> </div> <div> <h2>&nbsp;</h2> <h2>File Description</h2> </div> <div> <p>The dataset is provided in a csv file format, with each row representing a single message, including replies, quotes, and retweets. The file contains the following columns:&nbsp;&nbsp;</p> <p>'TweetID': Represents the tweet ID.&nbsp;&nbsp;&nbsp;</p> </div> <div> <p>'Username': Represents the username who published the tweet (if it is a retweet, it will be the user who retweetet the original tweet.&nbsp;&nbsp;&nbsp;&nbsp;</p> </div> <div> <p>'Text': Represents the full text of the tweet (not pre-processed).&nbsp;&nbsp;</p> </div> <div> <p>'CreateDate': Represents the date the tweet was created.&nbsp;&nbsp;&nbsp;&nbsp;</p> </div> <div> <p>'Biased': Represents the labeled by our annotators if the tweet is biased (1) or not (0).&nbsp;&nbsp;</p> </div> <div> <p>'Calling_Out': Represents the label by our annotators if the tweet is calling out bias against minority groups (1) or not (0).&nbsp;&nbsp;</p> </div> <div> <p>'Keyword': Represents the keyword that was used in the query. The keyword can be in the text, including mentioned names, or the username.&nbsp;&nbsp;&nbsp;&nbsp;</p> </div> <div> <p>&nbsp;&lsquo;Cohort&rsquo;: Represents the year the data was annotated (class of 2022 or class of 2023)&nbsp;</p> </div> <div> <h2>&nbsp;</h2> <h2>Acknowledgements&nbsp; &nbsp;</h2> </div> <div> <p>We are grateful for the technical collaboration with Indiana University's Observatory on Social Media (OSoMe). We thank all class participants for the annotations and contributions, including Kate Baba, Eleni Ballis, Garrett Banuelos, Savannah Benjamin, Luke Bianco, Zoe Bogan, Elisha S. Breton, Aidan Calderaro, Anaye Caldron, Olivia Cozzi, Daj Crisler, Jenna Eidson, Ella Fanning, Victoria Ford, Jess Gruettner, Ronan Hancock, Isabel Hawes, Brennan Hensler, Kyra Horton, Maxwell Idczak, Sanjana Iyer, Jacob Joffe, Katie Johnson, Allison Jones, Kassidy Keltner, Sophia Knoll, Jillian Kolesky, Emily Lowrey, Rachael Morara, Benjamin Nadolne, Rachel Neglia, Seungmin Oh, Kirsten Pecsenye, Sophia Perkovich, Joey Philpott, Katelin Ray, Kaleb Samuels, Chloe Sherman, Rachel Weber, Molly Winkeljohn, Ally Wolfgang, Rowan Wolke, Michael Wong, Jane Woods, Kaleb Woodworth, Aurora Young, Sydney Allen, Hundre Askie, Norah Bardol, Olivia Baren, Samuel Barth, Emma Bender, Noam Biron, Kendyl Bond, Graham Brumley, Kennedi Bruns, Leah Burger, Hannah Busche, Morgan Butrum-Griffith, Zoe Catlin, Angeli Cauley, Nathalya Chavez Medrano, Mia Cooper, Suhani Desai, Isabella Flick, Samantha Garcez, Isabella Grady, Macy Hutchinson, Sarah Kirkman, Ella Leitner, Elle Marquardt, Madison Moss, Ethan Nixdorf, Reya Patel, Mickey Racenstein, Kennedy Rehklau, Grace Roggeman, Jack Rossell, Madeline Rubin, Fernando Sanchez, Hayden Sawyer, Diego Scheker, Lily Schwecke, Brooke Scott, Megan Scott, Samantha Secchi, Jolie Segal, Katherine Smith, Constantine Stefanidis, Cami Stetler, Madisyn West, Alivia Yusefzadeh, Tayssir Aminou, Karen Fecht, Luciana Orrego-Hoyos, Hannah Pickett, and Sophia Tracy.&nbsp;</p> </div> <div> <p>This work used Jetstream2 at Indiana University through allocation HUM200003 from the Advanced Cyberinfrastructure Coordination Ecosystem: Services &amp; Support (ACCESS) program, which is supported by National Science Foundation grants #2138259, #2138286, #2138307, #2137603, and #2138296.&nbsp;</p> </div> <div> <p>&nbsp;</p> </div>

opencc-by-4.0Mar 2023View details →
zenodo40/100

Algerian Dialect Dataset Targeted Hate Speech, Offensive Language and Cyberbullying

<p>Algerian Dialect Dataset Targeted Hate Speech, Offensive Language and Cyberbullying.</p> <p>&nbsp;</p> <p>* To cite this dataset refer to&nbsp;<a href="http://dx.doi.org/10.12785/ijcds/130177" target="_blank" rel="nofollow noopener">http://dx.doi.org/10.12785/ijcds/130177</a><br>Mazari, A. C., &amp; Kheddar, H. (2023). "Deep Learning-based Analysis of Algerian Dialect Dataset Targeted Hate Speech, Offensive Language and Cyberbullying." IJCDS, 13(1).</p> <p>&nbsp;</p> <div> <p>* Due to the nature of this Dataset, comments contain offensiveness and hate speech. This does not reflect author values, however the aim is to providing a resource to help in detecting and preventing spread of such harmful content.</p> </div> <div> <h3>Features</h3> <ul> <li>Algerian Dialect</li> <li>Cyberbullying</li> <li>Hate speech</li> <li>Offensive Language</li> <li>Dialect Dataset</li> </ul> </div>

opencc-by-4.0Apr 2024View details →
zenodo40/100

On the Effectiveness of Text and Image Embeddings in Multimodal Hate Speech Detection

<p>Additional resources for the paper:</p> <h3><strong><a href="https://ieeexplore.ieee.org/abstract/document/10826088">On the Effectiveness of Text and Image Embeddings in Multimodal Hate Speech Detection.</a></strong></h3> <p>Lewis, N., Cavalcante, C. C., Boukouvalas, Z., &amp; Corizzo, R.</p> <p><em>2024 IEEE International Conference on Big Data (BigData)</em> (pp. 3277-3281). IEEE.</p> <pre>&nbsp;</pre> <p>&nbsp;</p> <p>MMHS150K [1] is a manually labeled multimodal dataset that contains $150000$ tweets with two modalities: text, and &nbsp;corresponding image. Tweets are collected from September 2018 until February 2019 and are labeled according to different types of hate speech: no attacks to any community, racist, sexist, homophobic, religion-based attacks, or attacks to other communities.&nbsp;</p> <p>We extract vector embeddings leveraging different text (BERT, OpenAI) and image (ResNet, PVT, ViT) modele backbones and assess their effectiveness in the hate speech detection task.</p> <p>&nbsp;</p> <h2>Citation:</h2> <pre>@inproceedings{lewis2024effectiveness, title={On the Effectiveness of Text and Image Embeddings in Multimodal Hate Speech Detection}, author={Lewis, Nora and Cavalcante, Charles C and Boukouvalas, Zois and Corizzo, Roberto}, booktitle={2024 IEEE International Conference on Big Data (BigData)}, pages={3277--3281}, year={2024}, organization={IEEE} }</pre>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Detecting weak and strong Islamophobic hate speech on social media

<p>Data, code and annotation guidelines for our publication, &#39;Detecting weak and strong Islamophobic hate speech on social media&#39; (2019).</p>

opencc-by-4.0Sep 2019View details →
zenodo40/100

Amharic Hate Speech Detection Dataset

<p>Amharic Hate Speech Detection Dataset V1</p> <p>To contribute for the research and development of hate speech detection in Amharic language from social media, we are glad to release our hate speech dataset we prepared from the Ethiopian Broadcasting Corporation (EBC) Facebook page (<a href="https://www.facebook.com/EBCzena">https://www.facebook.com/EBCzena</a>), and some chosen Facebook page (<a href="https://www.facebook.com/604407519910492">https://www.facebook.com/604407519910492</a>) that we found potential hateful comments.</p> <p>We extracted comments/posts pertaining to race, religion, and ethnicity using the Facepager API, resulting in a set of 30,000 comments between April 15, 2019 and December 15, 2019. A total of 5,000 comments/posts were chosen at random for annotation. Three annotators (two candidate PhD. in Linguistics and one MSc. in Law) manually annotated the selected samples as &ldquo;<strong>Hate</strong>&rdquo; or &ldquo;<strong>not</strong>-<strong>Hate</strong>&rdquo; resulting 2,000 (1000 hate and 1000 non-hate) labeled comments because of majority vote among the annotators.</p> <p>For the labeling procedure, the annotators used Ethiopian government&rsquo;s hate speech and misinformation prevention and suppression proclamation <a href="https://www.accessnow.org/cms/assets/uploads/2020/05/Hate-Speech-and-Disinformation-Prevention-and-Suppression-Proclamation.pdf">https://www.accessnow.org/cms/assets/uploads/2020/05/Hate-Speech-and-Disinformation-Prevention-and-Suppression-Proclamation.pdf</a>, as well as our definition of hate speech and the hate speech characterization lists proposed in Fino (2020) <a href="https://doi.org/10.1093/jicj/mqaa023">https://doi.org/10.1093/jicj/mqaa023</a>) were provided to the annotators.</p> <p>Accordingly, a speech is labeled as &ldquo;<strong>Hate</strong>&rdquo; when:</p> <ul> <li>&ldquo;the speech targets a group or individual as a member of a group (ethnicity, race, religion)&rdquo;</li> <li>&ldquo;the speech content in the message expresses hatred&rdquo;</li> <li>&ldquo;the speech causes a harm&rdquo;</li> <li>&ldquo;the speaker intends harm or bad activity&rdquo;</li> <li>&ldquo;the speech incites bad actions&rdquo;</li> <li>&ldquo;the speech is either public and directed at a member of the group&rdquo;</li> <li>&ldquo;the context makes violent response possible&rdquo;</li> </ul>

opencc-by-4.0Jun 2021View details →
zenodo36/100

Mpox Narrative on Instagram: A Labeled Multilingual Dataset of Instagram Posts on Mpox for Sentiment, Hate Speech, and Anxiety Analysis

<p><strong>Please cite the following paper when using this dataset</strong>:</p> <p>N. Thakur, &ldquo;Mpox narrative on Instagram: A labeled multilingual dataset of Instagram posts on mpox for sentiment, hate speech, and anxiety analysis,&rdquo; arXiv [cs.LG], 2024, URL: https://arxiv.org/abs/2409.05292</p> <p><strong>Abstract</strong></p> <p>The world is currently experiencing an outbreak of mpox, which has been declared a Public Health Emergency of International Concern by WHO. During recent virus outbreaks, social media platforms have played a crucial role in keeping the global population informed and updated regarding various aspects of the outbreaks. As a result, in the last few years, researchers from different disciplines have focused on the development of social media datasets focusing on different virus outbreaks. No prior work in this field has focused on the development of a dataset of Instagram posts about the mpox outbreak. The work presented in this paper (stated above) aims to address this research gap. It presents this <strong>multilingual dataset of</strong>&nbsp;<strong>60,127 Instagram posts</strong> about mpox, published between <strong>July 23, 2022, and September 5, 2024</strong>. This dataset contains Instagram posts about mpox in <strong>52 languages</strong>. For each of these posts, the Post ID, Post Description, Date of publication, language, and translated version of the post (translation to English was performed using the Google Translate API) are presented as separate attributes in the dataset.</p> <p>After developing this dataset, sentiment analysis, hate speech detection, and anxiety or stress detection were also performed. This process included classifying each post into</p> <ul> <li>one of the fine-grain sentiment classes, i.e., <strong>fear, surprise, joy, sadness, anger, disgust, or neutral</strong>,&nbsp;</li> <li><strong>hate or not hate</strong></li> <li><strong>anxiety/stress detected or no anxiety/stress detected</strong>.</li> </ul> <p>These results are presented as separate attributes in the dataset for the training and testing of machine learning algorithms for sentiment, hate speech, and anxiety or stress detection, as well as for other applications.&nbsp;</p> <p><strong>The 52 distinct languages in which Instagram posts are present in the dataset&nbsp;</strong><strong>are&nbsp;</strong>English, Portuguese, Indonesian, Spanish, Korean, French, Hindi, Finnish, Turkish, Italian, German, Tamil, Urdu, Thai, Arabic, Persian, Tagalog, Dutch, Catalan, Bengali, Marathi, Malayalam, Swahili, Afrikaans, Panjabi, Gujarati, Somali, Lithuanian, Norwegian, Estonian, Swedish, Telugu, Russian, Danish, Slovak, Japanese, Kannada, Polish, Vietnamese, Hebrew, Romanian, Nepali, Czech, Modern Greek, Albanian, Croatian, Slovenian, Bulgarian, Ukrainian, Welsh, Hungarian, and Latvian.&nbsp;</p> <p>The following table represents the data description for this dataset</p> <table> <tbody> <tr> <td> <p><strong>Attribute Name</strong></p> </td> <td> <p><strong>Attribute Description</strong></p> </td> </tr> <tr> <td> <p>Post ID</p> </td> <td> <p>Unique ID of each Instagram post</p> </td> </tr> <tr> <td> <p>Post Description</p> </td> <td> <p>Complete description of each post in the language in which it was originally published</p> </td> </tr> <tr> <td> <p>Date</p> </td> <td> <p>Date of publication in MM/DD/YYYY format</p> </td> </tr> <tr> <td> <p>Language</p> </td> <td> <p>Language of the post as detected using the Google Translate API</p> </td> </tr> <tr> <td> <p>Translated Post Description</p> </td> <td> <p>Translated version of the post description. All posts which were not in English were translated into English using the Google Translate API. No language translation was performed for English posts.</p> </td> </tr> <tr> <td> <p>Sentiment</p> </td> <td> <p>Results of sentiment analysis (using translated Post Description) where each post was classified into one of the sentiment classes: fear, surprise, joy, sadness, anger, disgust, and neutral</p> </td> </tr> <tr> <td> <p>Hate</p> </td> <td> <p>Results of hate speech detection (using translated Post Description) where each post was classified as hate or not hate</p> </td> </tr> <tr> <td> <p>Anxiety or Stress</p> </td> <td> <p>Results of anxiety or stress detection (using translated Post Description) where each post was classified as stress/anxiety detected or no stress/anxiety detected.</p> </td> </tr> </tbody> </table>

opencc-by-4.0Dec 2023View details →
zenodo36/100

Indonesian Foreign Policy towards Iran: Shia Hate speech on X social media with LSTM and SVM Analysis

<p>This table and figure are integral components of research on Shia hate speech on social media X, a critical issue in the context of identity politics in Indonesia and globally. This study is of paramount importance as it delves into the identity politics often exploited by politicians in Indonesia and around the world. The Shia community's support for President Jokowi in the first and second stages of the Election was met with hate speech from the opposition group. The study further investigates whether this Shia hate speech is linked to the government's policy towards Iran, a country known for its Shia ideology. The study is presented in three parts:<br>1. Table detailing the sentiment analysis process and results, which were conducted using advanced machine learning techniques such as SVM and LSTM. This approach significantly enhances the accuracy and reliability of the study's findings.<br>2. Figure in the form of a graph related to the study results and the results of comments from the Indonesian public about Shia.<br>3. This research is backed by a comprehensive dataset comprising public comments from Indonesia on Shia. This extensive data collection ensures the study's conclusions are thorough and reliable.</p> <p>4. Processing of machine learning</p>

opencc-by-4.0Oct 2024View details →
zenodo32/100

League of Legends and hate speech: a corpus for comments in Twitch.tv

<p>League of Legends (LOL) is the most popular game on PC, drawing 8 million concurrent players. A common activity of gamers, besides playing games, is to watch other players presenting tips and tricks. Streaming platforms allow some players to show gameplays and live games. <a href="https://www.twitch.tv/">Twitch.tv</a> is the world&acute;s leading live streaming platform.&nbsp;</p> <p>Considering that hate speech is a ubiquitous problem in online gaming, we collected &nbsp;985,766 comments from five videos of the top 10 &nbsp;LOL streamers in Twitch.tv platform.&nbsp;</p> <p>The dataset is freely available in a single file, ensembling all videos/players; and divided by players as well.&nbsp;</p> <p>These comments are a rich data source for opinion mining, sentiment analysis, topic modeling, and hate speech detection (including sexism and racism).</p> <ul> </ul>

opencc-by-4.0Mar 2020View details →
zenodo32/100

HaterNet a system for detecting and analyzing hate speech in Twitter

<p>This dataset consists of&nbsp; two corpuses used in the paper &quot;Detecting and analyzing hate speech in Twitter: HaterNet a system in the Spanish prevention of hate crime office&quot;. A first one based on tweets collected at different random dates between February 2017 and December 2017 with a final size of 2 million tweets. A second one with&nbsp;6,000 tweets labeled as described in the paper as hate containing or not.</p>

opencc-by-4.0Mar 2019View details →
zenodo32/100

HOCON34k: A Corpus of Hate speech in Online Comments from German Newspapers

<p>We have compiled a dataset containing 34,223 comments in German, authored by users from online-platforms associated with public discourse in German newspapers. Each comment was annotated for hate speech and the adequacy of contextual information by a group of 29 volunteers, using a binary annotation approach. The inter-rater reliability for hate speech is 0.4428 across all annotators and increases to 0.6078 when considering an optimized subset of 12 annotators, as measured by Fleiss&rsquo; Kappa. Additionally, we present a baseline text classification using BERT, achieving an MCC-score up to 0.32 and an F2-score up to 0.64 in our initial experiment on this new corpus. The data set, named HOCON34k, comprising German hate speech comments from newspapers, is publicly available for research purposes.</p>

opencc-by-4.0Dec 2024View details →
zenodo28/100

Hate Speech

Open the record for dataset details and reuse information.

opencc-by-4.0Jul 2024View details →
zenodo24/100

Supplementary material to 'Automatic Identification of Hate Speech – A Case-Study of Alt-Right YouTube Videos'

<p>The associated files have been created for and is analysed in a fortcoming article entitled&nbsp;<em>Automatic Identification of Hate Speech &ndash; A Case-Study of Alt-Right YouTube Videos'. </em>The material is divided into six tables as follows:</p> <table> <tbody> <tr> <td>Sentence top 5%</td> <td>The 19th 20-quantile predicted most hateful sentences</td> </tr> <tr> <td>Sentence bottom 5%</td> <td>The bottom 20-quantile predicted moste hatefull sentences (the least likely to contain hatespeech)</td> </tr> <tr> <td>Paragraphs</td> <td>Prediction and annotation of paragraphs</td> </tr> <tr> <td>Video top 10%</td> <td>Titles of the top decile predicted hateful videos</td> </tr> <tr> <td>Video bottom 10%</td> <td>Titles of the bottom decile predicted hateful videos</td> </tr> <tr> <td>Video bottom 10% - Alt right</td> <td>Titles of the bottom decile predicted hateful videos without History</td> </tr> </tbody> </table> <p>The data is uploaded in two formats:</p> <p><strong>Excel file:&nbsp;</strong>Automatic_Detection_of_Hate_Speech_a_Case-Study_of_Alt-Right_Videos.xlsx contains all six tables in one file, with a supplementary <em>codebook.&nbsp;</em></p> <p><strong>Tab Separated Values (TSV):</strong> Each file correspond to a single sheet from the excel file, and are named accordingly. UTF-8 Encoded.<strong><br></strong></p>

restrictedcc-by-4.0Jan 2024View details →
zenodo20/100

Profiling Hate Speech Spreaders on Twitter

<p><strong>Task</strong></p> <p>Hate speech (HS) is commonly defined as any communication that disparages a person or a group on the basis of some characteristic such as race, colour, ethnicity, gender, sexual orientation, nationality, religion, or other characteristics. Given the huge amount of user-generated contents on Twitter, the problem of detecting, and therefore possibly contrasting the HS diffusion, is becoming fundamental, for instance for fighting against misogyny and xenophobia. To this end, in this task, we aim at identifying possible hate speech spreaders on Twitter as a first step towards preventing hate speech from being propagated among online users.</p> <p>After having addressed several aspects of author profiling in social media from 2013 to 2020 (fake news spreaders, bot detection, age and gender, also together with personality, gender and language variety, and gender from a multimodality perspective), this year we aim at investigating if it is possible to discriminate authors that have shared some hate speech in the past from those that, to the best of our knowledge, have never done it.</p> <p>As in previous years, we propose the task from a&nbsp;<strong>multilingual</strong>&nbsp;perspective:</p> <ul> <li>English</li> <li>Spanish</li> </ul> <p><strong>NOTE:</strong>&nbsp;Although we recommend participating in both languages (English and Spanish), it is possible to address the problem just for one language.</p> <p>Award</p> <p>We are happy to announce that the best performing team at the 9th International Competition on Author Profiling will be awarded 300,- Euro sponsored by&nbsp;<a href="https://www.symanto.net/"><strong>Symanto</strong></a></p> <p><strong>Data</strong></p> <p><strong>Input</strong></p> <p>The uncompressed dataset consists of a folder per language (en, es). Each folder contains:</p> <ul> <li>An XML file per author (Twitter user) with 100 tweets. The name of the XML file corresponding to the unique author id.</li> <li>A truth.txt file with the list of authors and the ground truth.</li> </ul> <p>The format of the XML files is:</p> <pre> &lt;author lang=&quot;en&quot;&gt; &lt;documents&gt; &lt;document&gt;Tweet 1 textual contents&lt;/document&gt; &lt;document&gt;Tweet 2 textual contents&lt;/document&gt; ... &lt;/documents&gt; &lt;/author&gt; </pre> <p>The format of the truth.txt file is as follows. The first column corresponds to the author id. The second column contains the truth label.</p> <pre> b2d5748083d6fdffec6c2d68d4d4442d:::0 2bed15d46872169dc7deaf8d2b43a56:::0 8234ac5cca1aed3f9029277b2cb851b:::1 5ccd228e21485568016b4ee82deb0d28:::0 60d068f9cafb656431e62a6542de2dc0:::1 ... </pre> <p><strong>Output</strong></p> <p>Your software must take as input the absolute path to an unpacked dataset, and has to output for each document of the dataset a corresponding XML file that looks like this:</p> <pre> &lt;author id=&quot;author-id&quot; lang=&quot;en|es&quot; type=&quot;0|1&quot; /&gt; </pre> <p>The naming of the output files is up to you. However, we recommend using the author-id as filename and &quot;XML&quot; as an extension.</p> <p><strong>IMPORTANT!</strong>&nbsp;Languages should not be mixed. A folder should be created for each language and place inside only the files with the prediction for this language.</p> <p><strong>Evaluation</strong></p> <p>The performance of your system will be ranked by accuracy. For each language, we will calculate individual accuracies in discriminating between the two classes. Finally, we will average the accuracy values per language to obtain the final ranking.</p> <p><strong>Related Work</strong></p> <ul> <li>[1] Valerio Basile, Cristina Bosco, Elisabetta Fersini, Dora Nozza, Viviana Patti, Francisco Rangel, Paolo Rosso, Manuela Sanguinetti (2019).&nbsp;<a href="http://personales.upv.es/prosso/resources/BasileEtAl_SemEval19.pdf">SemEval-2019 task 5: Multilingual detection of hate speech against immigrants and women in Twitter.&nbsp;</a>Proc. SemEval 2019</li> <li>[2] Fabio Poletto, Valerio Basile, Manuela Sanguinetti, Cristina Bosco, Viviana Patti (2020).&nbsp;<a href="https://link.springer.com/article/10.1007/s10579-020-09502-8">Resources and benchmark corpora for hate speech detection: a systematic review.&nbsp;</a>Language Resources &amp; Evaluation. https://doi.org/10.1007/s10579-020-09502-8</li> <li>[3] Paula Fortuna, S&eacute;rgio Nunes (2018).&nbsp;<a href="https://dl.acm.org/doi/10.1145/3232676">A survey on automatic detection of hate speech in text.&nbsp;</a>ACM Computing Surveys (CSUR) 51.4</li> <li>[4] Maria Anzovino, Elisabetta Fersini, Paolo Rosso (2018).&nbsp;<a href="https://link.springer.com/chapter/10.1007/978-3-319-91947-8_6">Automatic Identification and Classification of Misogynistic Language on Twitter.&nbsp;</a>In: Proc. 23rd Int. Conf. on Applications of Natural Language to Information Systems, NLDB-2018, Springer-Verlag, LNCS(10859), pp. 57-64</li> <li>[5] Elisabetta Fersini, Paolo Rosso, Maria Anzovino (2018).&nbsp;<a href="http://personales.upv.es/prosso/resources/FersiniEtAl_IberEval18.pdf">Overview of the task on automatic misogyny identification at IberEval 2018.&nbsp;</a>Proc. IberEval 2018</li> <li>[6] Elisabetta Fersini, Dora Nozza, Paolo Rosso (2018).&nbsp;<a href="http://personales.upv.es/prosso/resources/FersiniEtAl_Evalita18.pdf">Overview of the Evalita 2018 task on automatic misogyny identification (AMI). Proc.&nbsp;</a>EVALITA 2018</li> <li>[7] Cristina Bosco, Felice Dell&#39;Orletta, Fabio Poletto, Manuela Sanguinetti, Maurizio Tesconi (2018).&nbsp;<a href="https://pdfs.semanticscholar.org/3eae/e4b2b8d9c7de52ba2386c73bb30097ec111c.pdf">Overview of the EVALITA 2018 hate speech detection task.&nbsp;</a>Proc. EVALITA 2018</li> <li>[8] Samuel Caetano da Silva, Thiago Castro Ferreira, Ricelli Moreira Silva Ramos, Ivandre Paraboni (2020).&nbsp;<a href="https://www.cys.cic.ipn.mx/ojs/index.php/CyS/article/view/3478">Data-driven and psycholinguistics motivated approaches to hate speech detection.&nbsp;</a>Computaci&oacute;n y Sistemas, 24(3): 1179&ndash;1188</li> <li>[9] Stiven Zimmerman, Udo Kruschwitz, Cris Fox (2018).&nbsp;<a href="https://www.aclweb.org/anthology/L18-1404.pdf">Improving hate speech detection with deep learning ensembles.&nbsp;</a>In Proc. of the Eleventh Int. Conf. on Language Resources and Evaluation (LREC 2018)</li> <li>[10] Francisco Rangel, Anastasia Giachanou, Bilal Ghanem, Paolo Rosso.&nbsp;<a href="http://ceur-ws.org/Vol-2696/paper_267.pdf">Overview of the 8th Author Profiling Task at PAN 2020: Profiling Fake News Spreaders on Twitter.&nbsp;</a>In: L. Cappellato, C. Eickhoff, N. Ferro, and A. N&eacute;v&eacute;ol (eds.) CLEF 2020 Labs and Workshops, Notebook Papers. CEUR Workshop Proceedings.CEUR-WS.org, vol. 2696</li> <li>[11] Francisco Rangel and Paolo Rosso.&nbsp;<a href="http://ceur-ws.org/Vol-2380/paper_263.pdf">Overview of the 7th Author Profiling Task at PAN 2019: Bots and Gender Profiling in Twitter.&nbsp;</a>In: L. Cappellato, N. Ferro, D. E. Losada and H. M&uuml;ller (eds.) CLEF 2019 Labs and Workshops, Notebook Papers. CEUR Workshop Proceedings.CEUR-WS.org, vol. 2380</li> <li>[12] Francisco Rangel, Paolo Rosso, Martin Potthast, Benno Stein.&nbsp;<a href="http://ceur-ws.org/Vol-2125/invited_paper_15.pdf">Overview of the 6th author profiling task at pan 2018: multimodal gender identification in Twitter.</a>&nbsp;In: CLEF 2018 Labs and Workshops, Notebook Papers. CEUR Workshop Proceedings. CEUR-WS.org, vol. 2125.</li> <li>[13] Francisco Rangel, Paolo Rosso, Martin Potthast, Benno Stein.&nbsp;<a href="http://ceur-ws.org/Vol-1866/invited_paper_11.pdf">Overview of the 5th Author Profiling Task at PAN 2017: Gender and Language Variety Identification in Twitter.</a>&nbsp;In: Cappellato L., Ferro N., Goeuriot L, Mandl T. (Eds.) CLEF 2017 Labs and Workshops, Notebook Papers. CEUR Workshop Proceedings. CEUR-WS.org, vol. 1866.</li> <li>[14] Francisco Rangel, Paolo Rosso, Ben Verhoeven, Walter Daelemans, Martin Pottast, Benno Stein.&nbsp;<a href="http://ceur-ws.org/Vol-1609/16090750.pdf">Overview of the 4th Author Profiling Task at PAN 2016: Cross-Genre Evaluations.</a>&nbsp;In: Balog K., Capellato L., Ferro N., Macdonald C. (Eds.) CLEF 2016 Labs and Workshops, Notebook Papers. CEUR Workshop Proceedings. CEUR-WS.org, vol. 1609, pp. 750-784</li> <li>[15] Francisco Rangel, Fabio Celli, Paolo Rosso, Martin Pottast, Benno Stein, Walter Daelemans.&nbsp;<a href="http://personales.upv.es/prosso/resources/RangelEtAl_PAN15.pdf%22">Overview of the 3rd Author Profiling Task at PAN 2015.</a>In: Linda Cappelato and Nicola Ferro and Gareth Jones and Eric San Juan (Eds.): CLEF 2015 Labs and Workshops, Notebook Papers, 8-11 September, Toulouse, France. CEUR Workshop Proceedings. ISSN 1613-0073, http://ceur-ws.org/Vol-1391/,2015.</li> <li>[16] Francisco Rangel, Paolo Rosso, Irina Chugur, Martin Potthast, Martin Trenkmann, Benno Stein, Ben Verhoeven, Walter Daelemans.&nbsp;<a href="http://ceur-ws.org/Vol-1180/CLEF2014wn-Pan-RangelEt2014.pdf">Overview of the 2nd Author Profiling Task at PAN 2014.</a>&nbsp;In: Cappellato L., Ferro N., Halvey M., Kraaij W. (Eds.) CLEF 2014 Labs and Workshops, Notebook Papers. CEUR-WS.org, vol. 1180, pp. 898-827.</li> <li>[17] Francisco Rangel, Paolo Rosso, Moshe Koppel, Efstatios Stamatatos, Giacomo Inches.&nbsp;<a href="http://ceur-ws.org/Vol-1179/CLEF2013wn-PAN-RangelEt2013.pdf">Overview of the Author Profiling Task at PAN 2013.</a>&nbsp;In: Forner P., Navigli R., Tufis D. (Eds.)Notebook Papers of CLEF 2013 LABs and Workshops. CEUR-WS.org, vol. 1179</li> <li>[18] Francisco Rangel and Paolo Rosso&nbsp;<a href="https://ojs.letras.up.pt/ojs/index.php/LLLD/article/download/6119/5761">On the Implications of the General Data Protection Regulation on the Organisation of Evaluation Tasks.&nbsp;</a>In: Language and Law / Linguagem e Direito, Vol. 5(2), pp. 80-102</li> <li>[19] Francisco Rangel, Marc Franco-Salvador, Paolo Rosso&nbsp;<a href="https://arxiv.org/abs/1705.10754">A Low Dimensionality Representation for Language Variety Identification.&nbsp;</a>In: Postproc. 17th Int. Conf. on Comput. Linguistics and Intelligent Text Processing, CICLing-2016, Springer-Verlag, Revised Selected Papers, Part II, LNCS(9624), pp. 156-169 (arXiv:1705.10754)</li> </ul>

restrictedMar 2021View details →
zenodo16/100

Datasets for "Auditing Elon Musk's Impact on Hate Speech and Bots"

<p>Datasets for the publication "Auditing Elon Musk's Impact on Hate Speech and Bots" [1].</p> <p><strong>File information:</strong></p> <ul> <li>baseline_tweet_ids_2022.csv, hate_tweet_ids_2022.csv: List of IDs and their corresponding dates from the "baseline" and "hate" samples of tweets used in the publication, respectively.&nbsp;The former is used to create the number of baseline tweets each day (&lsquo;baseline_freq.csv&rsquo;) while the latter is used to create the number of hate tweets each day (&lsquo;hate_freq.csv'). We share the date a tweet was made as well as its tweet ID from which you can find the original tweet&rsquo;s URL <a href="https://developer.twitter.com/en/blog/community/2020/getting-to-the-canonical-url-for-a-post">with the help of this web page</a>.&nbsp; As you explore these data, you may notice in a minority of cases hate tweets that are not hateful or, alternatively, baseline tweets that are hateful. This is a product of our filtering method used to collect and analyze tweets at scale. We always look forward to hearing your suggestions to improve the tweet filtering process.</li> <li>baseline_freq.csv, hate_freq.csv: Number of collected tweets per day for the baseline and hate samples, respectively. The file 'freq_data.py' is used to calculate these frequencies from the raw data. Feel free to consult this if you have questions about how the frequencies are calculated (or if you want to change how the data are aggregated).&nbsp;<strong>Use these to recreate Figure 2 from Hickey et al [1].</strong></li> <li>user_hate_levels_per_day.csv: CSV file with dates (YYYY-MM-DD format) and the mean proportion of slurs used by hateful users each day from October 1st to November 30th, 2022.&nbsp;<strong>Use these data to recreate Figure 1 from Hickey et al [1].&nbsp;</strong>See the Methods section of Hickey et al [1]. for details.</li> <li>hate_keywords.txt: Words used to query the Twitter Academic API for hate tweets.</li> <li>unfiltered_tweets_containing_hate_words.csv: All tweets with hate words collected with values for Perspective API attributes.</li> </ul> <p><strong>Reference:</strong></p> <p>1. Hickey, D., Schmitz, M., Fessler, D.M.T, Smaldino, P., Muric, G., &amp; Burghardt, K. Auditing Elon Musk's Impact on Hate Speech and Bots. In Proceedings of the 17th International AAAI Conference on Web and Social Media, (2023).</p>

restrictedcc-by-4.0Dec 2023View details →
zenodo16/100

SWAHILI AND CODE-SWITCHED ENGLISH-SWAHILI POLITICAL HATE SPEECH DETECTION TEXTUAL DATASET

<p>This dataset consist of swahili and code switched English-Swahili tweets labeled for hate speech type,the target and language.</p> <p>The dataset can be accessed upon request from the authors via email addresses:</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;-nodhianbo@gmail.com</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;-endeshanelly@gmail.com</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;-mscci00083@student.maseno.ac.ke</p>

restrictedcc-by-4.0Nov 2024View details →
zenodo16/100

Hateful Messages: A Conversational Data Set of Hate Speech produced by Adolescents on Discord

<p>With the rise of social media, a rise of hateful content can be observed. Even though the understanding and definitions of hate speech varies, platforms, communities, and legislature all acknowledge the problem. Therefore, adolescents are a new and active group of social media users. The majority of adolescents experience or witness online hate speech. Research in the field of automated hate speech classification has been on the rise and focuses on aspects such as bias, generalizability, and performance. To increase generalizability and performance, it is important to understand biases within the data. This research addresses the bias of youth language within hate speech classification and contributes by providing a modern and anonymized hate speech youth language data set consisting of 88.395 annotated chat messages. The data set consists of publicly available online messages from the chat platform Discord. ~6,42\% of the messages were classified by a self-developed annotation schema as hate speech. For 35.553 messages, the user profiles provided age annotations setting the average author age to under 20 years old.</p>

restrictedMay 2023View details →
zenodo12/100

Data for Reported user-generated online hate speech: The 'ecosystem', frames, and ideologies

<p>This is the dataset for the article entitled Reported user-generated online hate speech: The &#39;ecosystem&#39;, frames, and ideologies. The same dataset is provided in two different formats: comma-separated values (.csv) and Excel format (.xlsx). A basic legend to the data is provided separately in the corresponding PDF document. An extended legend to the data is available at:&nbsp;<strong><a href="https://doi.org/10.5281/zenodo.6656185">https://doi.org/10.5281/zenodo.6656185</a></strong>.</p>

restrictedJun 2022View details →
zenodo12/100

Hate speech and personal attack dataset in French social media

<p>This dataset contains 29109&nbsp;French&nbsp;tweet ids and corresponding annotations for Hate Speech label&nbsp;and 39109&nbsp;French&nbsp;tweet ids and corresponding annotations for Personal attack&nbsp;label.&nbsp;The creation of this dataset was part of the project DACHS &ldquo;A Data-driven Approach to Countering Hate Speech&rdquo; funded by the Rights, Equality and Citizenship Programme of the European Union.</p>

restrictedOct 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record